Written and maintained by CASRAI Editorial Board
Last updated
Biostatistics is the application of statistics to medicine, public health, and biology: designing studies, analyzing the data they produce, and stating how much confidence the results justify. It is what turns a clinical or scientific question — does this treatment work, does this exposure raise risk, how accurate is this test — into evidence that can be defended. This guide covers the definition, how biostatistics differs from statistics and epidemiology, its core methods, what biostatisticians do, the software they use, how regulators fit in, and how to train for the field.
Biostatistics is the branch of statistics concerned with the design and analysis of studies in medicine, public health, and biology — from a single laboratory experiment to a multi-country clinical trial. It is one of the major scientific disciplines covered in CASRAI’s guide to the branches of science; this page goes deeper on what biostatistics actually studies, how it is organized as a field, who funds it, and how someone trains to work in it.
What Is Biostatistics?
Biostatistics is the application and development of statistical methods to answer questions arising in biology, medicine, and public health. Where general statistics develops methods that could in principle be applied to any field with numeric data, biostatistics is defined by its subject matter: living systems, and the specific data-generating problems those systems create — small and expensive samples, measurements taken repeatedly on the same subject over time, outcomes that occur (or don’t) rather than being measured on a continuous scale, and interventions that cannot ethically be assigned the way an engineer might assign a stress test. A biostatistician’s core job is to translate a scientific or clinical question — does this drug extend survival, does this exposure raise disease risk, is this diagnostic test accurate enough to use — into a study design and analytical plan that can actually answer it with a defensible degree of confidence, then to carry out that analysis and interpret what it does and does not show.
The field’s questions cluster around a few recurring problems: how to design a study (how many subjects are needed, how should they be assigned to groups, what should be measured and when) so that it has a realistic chance of detecting a real effect if one exists; how to handle the fact that biological data is rarely as clean as an idealized statistical model assumes — missing follow-up visits, participants who drop out non-randomly, measurements clustered within the same patient, hospital, or family; how to distinguish a genuine effect from noise, confounding, or multiple-comparison chance; and how to communicate the resulting uncertainty (a confidence interval, a hazard ratio, a posterior probability) honestly rather than overstating certainty a study’s design cannot support.
Biostatistics vs. statistics, epidemiology, and bioinformatics
These four terms are often used loosely as synonyms, but they name genuinely distinct — if overlapping — activities:
- Statistics is the parent discipline: the general theory of collecting, analyzing, and drawing inference from data, developed and taught independently of any single application domain.
- Biostatistics is statistics specifically applied to and shaped by biological, medical, and public-health data — the methods above (survival analysis, longitudinal models, clinical trial design) exist largely because biomedical data has recurring structural features that generic statistical methods handle poorly.
- Epidemiology is a substantive field — the study of how disease and health outcomes are distributed in populations and what causes that distribution — that relies heavily on biostatistical methods but is not itself a branch of statistics; an epidemiologist designs the study question and interprets it in a population-health context, often working alongside a biostatistician who designs the sampling and analysis. CASRAI’s epidemiology guide covers that discipline in full.
- Bioinformatics sits at the intersection of biostatistics, computer science, and molecular biology, focused on developing computational methods and software to manage and analyze large-scale biological data, especially genomic and proteomic data. Statistical genetics — described below — is the part of biostatistics that overlaps most directly with bioinformatics.
For a side-by-side treatment of the first two, see biostatistics vs. epidemiology; for how statistics relates to the newer label of data science, see data science vs. statistics.
Core Methods at a Glance
Whatever the sub-discipline, most biostatistical work draws on a handful of method families:
- Study design — choosing the comparison, the sample size, the randomization or sampling scheme, and the measurements before data are collected. Design decisions cannot be repaired afterward by clever analysis.
- Regression — linear, logistic, and Poisson models (and generalized linear models more broadly) to estimate how an outcome relates to a treatment or exposure while adjusting for other variables.
- Survival analysis — methods such as Kaplan–Meier curves and Cox proportional hazards models for time-to-event outcomes with censoring. See how to interpret a hazard ratio.
- Longitudinal and mixed models — for repeated measurements and clustered data; see mixed-effects models.
- Clinical trial design and interim analysis — randomization, power, group-sequential monitoring that allows a trial to stop early for benefit, harm, or futility, and adaptive designs. The phases these trials move through are explained in clinical trial phases.
- Bayesian methods — combining prior information with trial or study data, increasingly used where samples are small or designs adapt as data accrue.
What Biostatisticians Do
In clinical trials, a biostatistician typically helps write the protocol, calculates sample size, specifies the randomization, writes the statistical analysis plan before the data are unblinded, supports the independent committee that monitors accumulating safety and efficacy data, runs the final analysis, and contributes to the clinical study report and regulatory submission. See the statistical analysis plan template for what that pre-specification contains.
In public health, biostatisticians analyze surveillance and survey data, design and analyze cohort and case-control studies, estimate disease burden and trends, and evaluate whether programs and screening tests achieve their intended effect — usually alongside epidemiologists and other public-health practitioners.
Regulatory Touchpoints
Drug and device development is one of the most regulated settings in which biostatistics is practiced. The international guideline most often cited for the statistical side is ICH E9, Statistical Principles for Clinical Trials, issued by the International Council for Harmonisation and published by the U.S. FDA as guidance in September 1998. It was later supplemented by E9(R1), an addendum on estimands and sensitivity analysis, which the FDA published in May 2021. The addendum asks trial teams to state precisely which treatment effect a trial is meant to estimate — the estimand framework — and to test how robust the conclusions are to assumptions. Biostatisticians at sponsors, contract research organizations, and the FDA all work within this framework. Sponsors should always check the current version of any guidance directly with the regulator.
Major Sub-Disciplines Within Biostatistics
Biostatistics is not a single toolkit; practitioners typically specialize in one or two of the following areas, each shaped by a distinct kind of data and question.
Clinical trial biostatistics
Designs and analyzes controlled trials of medical interventions — determining sample size and statistical power, randomization scheme, interim monitoring rules, and the pre-specified analysis that will determine whether a treatment effect is declared. This is the largest single employer of biostatisticians in industry and academic medical centers, and the area most tightly regulated (see the funding and regulatory context below).
Survival and time-to-event analysis
Develops and applies methods for outcomes defined by when an event happens rather than whether it happens — time to death, disease progression, or relapse — where standard regression methods break down because follow-up time differs across subjects and some subjects never experience the event during the study (censoring). CASRAI’s guide on interpreting a hazard ratio covers the core output of this subfield in depth.
Statistical genetics and genomics
Applies statistical methods to genetic and genomic data — linking DNA variation to disease risk or trait variation, correcting for the enormous multiple-testing burden created by testing millions of genetic markers at once, and modeling the shared ancestry that makes individuals in a genetic study non-independent. This subfield overlaps heavily with bioinformatics and, on the trait side, with genetics as a discipline.
Biostatistics for epidemiology and public health
Develops and applies the methods used to study disease patterns in populations rather than controlled experiments — cohort and case-control study design, adjustment for confounding, and methods for surveillance data. CASRAI’s guide to the cohort study covers one of this subfield’s foundational designs.
Longitudinal and multilevel modeling
Handles data collected repeatedly on the same subjects, or nested within a shared structure (patients within hospitals, students within schools) — situations where observations are correlated rather than independent, which violates the assumptions of standard regression. See CASRAI’s guide on mixed-effects models for how this is handled in practice.
Bayesian biostatistics
Applies Bayesian inference — updating a prior probability with observed data to produce a posterior probability — to biomedical questions, increasingly used in adaptive clinical trial designs, rare-disease research (where sample sizes are too small for standard frequentist power), and diagnostic-test evaluation.
Diagnostic and screening test evaluation
Develops the statistical methods used to judge how well a diagnostic or screening test performs — sensitivity, specificity, and the tradeoffs between them. CASRAI’s guide to Youden’s J Index and ROC threshold selection covers one of this subfield’s standard tools.
Who Funds Biostatistics Research
Biostatistics research is funded largely as a component of the larger biomedical, public-health, and basic-science research it supports, rather than through a single dedicated funding stream — a pattern worth understanding on its own, since it shapes how biostatistics careers and grant applications are actually structured.
In the United States, several institutes within the National Institutes of Health (NIH) fund biostatistics research as an integral part of the clinical trials, cohort studies, and basic biomedical science they support — biostatistical methods development is typically funded either as a component of a larger disease-focused or mechanism-focused grant (a multi-project grant such as an NIH P01 program project grant will commonly include a dedicated biostatistics core), or through methods-focused awards from institutes whose mission spans basic quantitative and computational biomedical science. Outside NIH, the National Science Foundation funds core statistical methodology — including methods with direct biostatistical application — primarily through the Statistics program housed in its Division of Mathematical Sciences.
Beyond U.S. federal funding, major biomedical research foundations with a substantial global-health and population-health footprint — including the Wellcome Trust and the Bill & Melinda Gates Foundation — fund biostatistical and epidemiological methods work as part of their broader investments in clinical trials, cohort studies, and global-health research infrastructure, though (like NIH) rarely as a dedicated “biostatistics” funding line; the funding is embedded in the substantive research it supports. Readers should treat any specific, named grant program as something to verify directly against the funder’s current site before citing it — funding program names and structures change more often than the underlying research areas they support.
Methods, Tools, and Data Biostatisticians Work With
Day-to-day biostatistical work spans a fairly consistent toolkit across sub-disciplines:
- Study design tools — power and sample-size calculation, randomization schemes (simple, block, stratified), and the broader design frameworks covered in CASRAI’s statistical analysis plan guide, which documents how a trial’s analysis is pre-specified before data collection begins.
- Statistical software — R and Python (both open-source, extensible, and dominant in academic biostatistics), SAS (still the standard in much of the pharmaceutical industry and FDA submissions due to its validated, auditable output), and Stata (common in epidemiology and health-services research). CASRAI compares these directly in SAS vs. SPSS and R vs. Stata.
- Data capture and management systems — electronic data capture platforms used to collect clean, auditable trial and study data; CASRAI’s Qualtrics vs. REDCap comparison covers two of the most common academic tools in this category.
- Core inferential concepts — the p-value and its correct interpretation, and the more recently emphasized estimand framework, which formalizes exactly what quantity a clinical trial analysis is meant to estimate before deciding how to estimate it.
- Computing infrastructure — increasingly, high-performance and cloud computing for genomic and large-cohort analyses that exceed what a single workstation can process, particularly in statistical genetics and large-scale public-health surveillance work.
Careers and Training in Biostatistics
Biostatistics is typically taught as a graduate discipline. Most practicing biostatisticians hold a master’s degree (MS or MPH with a biostatistics concentration) or a PhD, usually earned in a biostatistics department housed within a school of public health, a medical school, or occasionally a standalone statistics department with a biomedical track. A master’s-level biostatistician commonly works as an applied analyst — implementing an analysis plan designed by a more senior biostatistician or collaborating directly on a defined study — while a PhD is generally expected for roles that involve developing new statistical methods, leading a trial’s statistical design independently, or serving as principal investigator on a methods-focused grant. Typical graduate coursework covers probability theory and mathematical statistics, generalized linear models, survival analysis, longitudinal/multilevel data analysis, clinical trial design, and increasingly, statistical computing and, for genetics-focused tracks, statistical genetics.
The field’s principal professional societies include the American Statistical Association (ASA), the general professional body for statisticians in the United States, which has sections and communities specifically focused on biopharmaceutical statistics and health-policy statistics; and the International Biometric Society, an international professional society specifically for statisticians and quantitative scientists working in the biosciences, whose North American regions (ENAR and WNAR) run some of the field’s most attended annual conferences.
The International Biometric Society describes itself as a society of regions serving statisticians, mathematicians, biological scientists, and others working on the collection and interpretation of data in the biosciences, and publishes the journal Biometrics. ENAR, its Eastern North American Region, holds an annual spring meeting. The American Statistical Association, founded in Boston in 1839, is the second-oldest continuously operating professional association in the United States.
For labor-market context, the U.S. Bureau of Labor Statistics reports for statisticians as an occupation (not biostatisticians specifically) a median annual wage of $105,650 in May 2025 and projected employment growth of 10 percent from 2025 to 2035, and notes that statisticians typically need a master’s degree, though some entry-level positions accept a bachelor’s. Biostatistics-specific figures vary by employer and sector.
Biostatisticians work across academic medical centers and schools of public health, pharmaceutical and biotechnology companies, contract research organizations (CROs) that run trials on sponsors’ behalf, and government agencies — including the FDA, which employs biostatisticians to review the statistical evidence behind drug and device applications, and public-health agencies that run population surveillance and outcomes research.
Frequently Asked Questions
Is biostatistics the same as epidemiology?
No. Epidemiology is the substantive study of how disease is distributed in populations and what causes that distribution; biostatistics is the statistical methodology epidemiologists rely on to design and analyze those studies. In practice the two fields work closely together and their graduate training overlaps significantly, but they are distinct disciplines with distinct professional identities.
What is the difference between biostatistics and statistics?
Statistics is the general theory of data collection and inference, developed independently of any one application area. Biostatistics is that theory applied to — and substantially shaped by — the recurring data problems of biology and medicine: censored survival data, repeated measurements on the same patient, small and expensive clinical samples, and correlated data within families, hospitals, or genetic pedigrees.
What degree do you need to become a biostatistician?
Most practicing biostatisticians hold at least a master’s degree in biostatistics, statistics, or a closely related quantitative field, typically earned within a school of public health or medical school. Roles that involve developing new statistical methods or leading a trial’s statistical design independently generally require a PhD.
What software do biostatisticians use?
R and Python dominate academic biostatistics; SAS remains the standard for much pharmaceutical-industry and regulatory (FDA submission) work because of its long track record of validated, auditable output; Stata is common in epidemiology and health-services research.
Who funds biostatistics research?
In the U.S., biostatistics research is funded largely as a component of the larger NIH-funded clinical trials, cohort studies, and basic biomedical research it supports, plus dedicated statistical-methodology funding from the National Science Foundation’s Division of Mathematical Sciences. Several major private foundations active in global and population health also fund biostatistical methods work as part of broader research investments, though rarely through a dedicated biostatistics-only program.
Do biostatisticians need to know how to program?
Yes. Day-to-day work is done in statistical software, most commonly R, SAS, or Stata, so programming in at least one of them is a core skill. Which one depends on the employer: pharmaceutical and regulatory work leans on SAS and R, while epidemiology and health-services research often use Stata or R.
Is there a regulatory standard for statistics in clinical trials?
The international reference point is ICH E9, Statistical Principles for Clinical Trials, with its E9(R1) addendum on estimands and sensitivity analysis. Regulators such as the FDA publish these as guidance, and the current text should be checked at the source.
How much do statisticians earn?
The U.S. Bureau of Labor Statistics reports a median annual wage of $105,650 for statisticians as of May 2025. That figure covers the whole occupation rather than biostatistics alone.
Related Fields
Biostatistics sits among several neighboring disciplines. Start with the broader statistics foundation, then see how it connects to epidemiology (and the epidemiology vs. public health distinction), public health, health informatics, bioinformatics, data science, and machine learning. For how trials are staged, read clinical trial phases.








