Written and maintained by CASRAI Editorial Board
Last updated
Biostatistics is the branch of statistics concerned with the design and analysis of studies in medicine, public health, and biology — from a single laboratory experiment to a multi-country clinical trial. It is one of the major scientific disciplines covered in CASRAI’s guide to the branches of science; this page goes deeper on what biostatistics actually studies, how it is organized as a field, who funds it, and how someone trains to work in it.
What Is Biostatistics?
Biostatistics is the application and development of statistical methods to answer questions arising in biology, medicine, and public health. Where general statistics develops methods that could in principle be applied to any field with numeric data, biostatistics is defined by its subject matter: living systems, and the specific data-generating problems those systems create — small and expensive samples, measurements taken repeatedly on the same subject over time, outcomes that occur (or don’t) rather than being measured on a continuous scale, and interventions that cannot ethically be assigned the way an engineer might assign a stress test. A biostatistician’s core job is to translate a scientific or clinical question — does this drug extend survival, does this exposure raise disease risk, is this diagnostic test accurate enough to use — into a study design and analytical plan that can actually answer it with a defensible degree of confidence, then to carry out that analysis and interpret what it does and does not show.
The field’s questions cluster around a few recurring problems: how to design a study (how many subjects are needed, how should they be assigned to groups, what should be measured and when) so that it has a realistic chance of detecting a real effect if one exists; how to handle the fact that biological data is rarely as clean as an idealized statistical model assumes — missing follow-up visits, participants who drop out non-randomly, measurements clustered within the same patient, hospital, or family; how to distinguish a genuine effect from noise, confounding, or multiple-comparison chance; and how to communicate the resulting uncertainty (a confidence interval, a hazard ratio, a posterior probability) honestly rather than overstating certainty a study’s design cannot support.
Biostatistics vs. statistics, epidemiology, and bioinformatics
These four terms are often used loosely as synonyms, but they name genuinely distinct — if overlapping — activities:
- Statistics is the parent discipline: the general theory of collecting, analyzing, and drawing inference from data, developed and taught independently of any single application domain.
- Biostatistics is statistics specifically applied to and shaped by biological, medical, and public-health data — the methods above (survival analysis, longitudinal models, clinical trial design) exist largely because biomedical data has recurring structural features that generic statistical methods handle poorly.
- Epidemiology is a substantive field — the study of how disease and health outcomes are distributed in populations and what causes that distribution — that relies heavily on biostatistical methods but is not itself a branch of statistics; an epidemiologist designs the study question and interprets it in a population-health context, often working alongside a biostatistician who designs the sampling and analysis. CASRAI’s epidemiology guide covers that discipline in full.
- Bioinformatics sits at the intersection of biostatistics, computer science, and molecular biology, focused on developing computational methods and software to manage and analyze large-scale biological data, especially genomic and proteomic data. Statistical genetics — described below — is the part of biostatistics that overlaps most directly with bioinformatics.
Major Sub-Disciplines Within Biostatistics
Biostatistics is not a single toolkit; practitioners typically specialize in one or two of the following areas, each shaped by a distinct kind of data and question.
Clinical trial biostatistics
Designs and analyzes controlled trials of medical interventions — determining sample size and statistical power, randomization scheme, interim monitoring rules, and the pre-specified analysis that will determine whether a treatment effect is declared. This is the largest single employer of biostatisticians in industry and academic medical centers, and the area most tightly regulated (see the funding and regulatory context below).
Survival and time-to-event analysis
Develops and applies methods for outcomes defined by when an event happens rather than whether it happens — time to death, disease progression, or relapse — where standard regression methods break down because follow-up time differs across subjects and some subjects never experience the event during the study (censoring). CASRAI’s guide on interpreting a hazard ratio covers the core output of this subfield in depth.
Statistical genetics and genomics
Applies statistical methods to genetic and genomic data — linking DNA variation to disease risk or trait variation, correcting for the enormous multiple-testing burden created by testing millions of genetic markers at once, and modeling the shared ancestry that makes individuals in a genetic study non-independent. This subfield overlaps heavily with bioinformatics and, on the trait side, with genetics as a discipline.
Biostatistics for epidemiology and public health
Develops and applies the methods used to study disease patterns in populations rather than controlled experiments — cohort and case-control study design, adjustment for confounding, and methods for surveillance data. CASRAI’s guide to the cohort study covers one of this subfield’s foundational designs.
Longitudinal and multilevel modeling
Handles data collected repeatedly on the same subjects, or nested within a shared structure (patients within hospitals, students within schools) — situations where observations are correlated rather than independent, which violates the assumptions of standard regression. See CASRAI’s guide on mixed-effects models for how this is handled in practice.
Bayesian biostatistics
Applies Bayesian inference — updating a prior probability with observed data to produce a posterior probability — to biomedical questions, increasingly used in adaptive clinical trial designs, rare-disease research (where sample sizes are too small for standard frequentist power), and diagnostic-test evaluation.
Diagnostic and screening test evaluation
Develops the statistical methods used to judge how well a diagnostic or screening test performs — sensitivity, specificity, and the tradeoffs between them. CASRAI’s guide to Youden’s J Index and ROC threshold selection covers one of this subfield’s standard tools.
Who Funds Biostatistics Research
Biostatistics research is funded largely as a component of the larger biomedical, public-health, and basic-science research it supports, rather than through a single dedicated funding stream — a pattern worth understanding on its own, since it shapes how biostatistics careers and grant applications are actually structured.
In the United States, several institutes within the National Institutes of Health (NIH) fund biostatistics research as an integral part of the clinical trials, cohort studies, and basic biomedical science they support — biostatistical methods development is typically funded either as a component of a larger disease-focused or mechanism-focused grant (a multi-project grant such as an NIH P01 program project grant will commonly include a dedicated biostatistics core), or through methods-focused awards from institutes whose mission spans basic quantitative and computational biomedical science. Outside NIH, the National Science Foundation funds core statistical methodology — including methods with direct biostatistical application — primarily through the Statistics program housed in its Division of Mathematical Sciences.
Beyond U.S. federal funding, major biomedical research foundations with a substantial global-health and population-health footprint — including the Wellcome Trust and the Bill & Melinda Gates Foundation — fund biostatistical and epidemiological methods work as part of their broader investments in clinical trials, cohort studies, and global-health research infrastructure, though (like NIH) rarely as a dedicated “biostatistics” funding line; the funding is embedded in the substantive research it supports. Readers should treat any specific, named grant program as something to verify directly against the funder’s current site before citing it — funding program names and structures change more often than the underlying research areas they support.
Methods, Tools, and Data Biostatisticians Work With
Day-to-day biostatistical work spans a fairly consistent toolkit across sub-disciplines:
- Study design tools — power and sample-size calculation, randomization schemes (simple, block, stratified), and the broader design frameworks covered in CASRAI’s statistical analysis plan guide, which documents how a trial’s analysis is pre-specified before data collection begins.
- Statistical software — R and Python (both open-source, extensible, and dominant in academic biostatistics), SAS (still the standard in much of the pharmaceutical industry and FDA submissions due to its validated, auditable output), and Stata (common in epidemiology and health-services research). CASRAI compares these directly in SAS vs. SPSS and R vs. Stata.
- Data capture and management systems — electronic data capture platforms used to collect clean, auditable trial and study data; CASRAI’s Qualtrics vs. REDCap comparison covers two of the most common academic tools in this category.
- Core inferential concepts — the p-value and its correct interpretation, and the more recently emphasized estimand framework, which formalizes exactly what quantity a clinical trial analysis is meant to estimate before deciding how to estimate it.
- Computing infrastructure — increasingly, high-performance and cloud computing for genomic and large-cohort analyses that exceed what a single workstation can process, particularly in statistical genetics and large-scale public-health surveillance work.
Careers and Training in Biostatistics
Biostatistics is typically taught as a graduate discipline. Most practicing biostatisticians hold a master’s degree (MS or MPH with a biostatistics concentration) or a PhD, usually earned in a biostatistics department housed within a school of public health, a medical school, or occasionally a standalone statistics department with a biomedical track. A master’s-level biostatistician commonly works as an applied analyst — implementing an analysis plan designed by a more senior biostatistician or collaborating directly on a defined study — while a PhD is generally expected for roles that involve developing new statistical methods, leading a trial’s statistical design independently, or serving as principal investigator on a methods-focused grant. Typical graduate coursework covers probability theory and mathematical statistics, generalized linear models, survival analysis, longitudinal/multilevel data analysis, clinical trial design, and increasingly, statistical computing and, for genetics-focused tracks, statistical genetics.
The field’s principal professional societies include the American Statistical Association (ASA), the general professional body for statisticians in the United States, which has sections and communities specifically focused on biopharmaceutical statistics and health-policy statistics; and the International Biometric Society, an international professional society specifically for statisticians and quantitative scientists working in the biosciences, whose North American regions (ENAR and WNAR) run some of the field’s most attended annual conferences.
Biostatisticians work across academic medical centers and schools of public health, pharmaceutical and biotechnology companies, contract research organizations (CROs) that run trials on sponsors’ behalf, and government agencies — including the FDA, which employs biostatisticians to review the statistical evidence behind drug and device applications, and public-health agencies that run population surveillance and outcomes research.
Frequently Asked Questions
Is biostatistics the same as epidemiology?
No. Epidemiology is the substantive study of how disease is distributed in populations and what causes that distribution; biostatistics is the statistical methodology epidemiologists rely on to design and analyze those studies. In practice the two fields work closely together and their graduate training overlaps significantly, but they are distinct disciplines with distinct professional identities.
What is the difference between biostatistics and statistics?
Statistics is the general theory of data collection and inference, developed independently of any one application area. Biostatistics is that theory applied to — and substantially shaped by — the recurring data problems of biology and medicine: censored survival data, repeated measurements on the same patient, small and expensive clinical samples, and correlated data within families, hospitals, or genetic pedigrees.
What degree do you need to become a biostatistician?
Most practicing biostatisticians hold at least a master’s degree in biostatistics, statistics, or a closely related quantitative field, typically earned within a school of public health or medical school. Roles that involve developing new statistical methods or leading a trial’s statistical design independently generally require a PhD.
What software do biostatisticians use?
R and Python dominate academic biostatistics; SAS remains the standard for much pharmaceutical-industry and regulatory (FDA submission) work because of its long track record of validated, auditable output; Stata is common in epidemiology and health-services research.
Who funds biostatistics research?
In the U.S., biostatistics research is funded largely as a component of the larger NIH-funded clinical trials, cohort studies, and basic biomedical research it supports, plus dedicated statistical-methodology funding from the National Science Foundation’s Division of Mathematical Sciences. Several major private foundations active in global and population health also fund biostatistical methods work as part of broader research investments, though rarely through a dedicated biostatistics-only program.








