Written and maintained by CASRAI Editorial Board
Last updated
Statistics is the discipline that studies how to collect, describe, analyze, and draw
defensible conclusions from data in the presence of uncertainty and variability. Wherever a
researcher observes only a sample rather than an entire population — a clinical trial with
a few hundred patients, a survey of a few thousand voters, a set of sensor readings from a
handful of field sites — statistics supplies the formal machinery for deciding what that
sample can and cannot responsibly tell you about the larger reality it was drawn from. It is
simultaneously a branch of mathematics (built on probability theory) and an applied,
problem-driven discipline that sits underneath nearly every other quantitative field of
research.
This guide answers “what is statistics” in real depth: its core questions and
methods, its major subfields, and — because this page is published by
CASRAI, a research-administration standards body — the
funding landscape, research methods, and career pathways that matter to anyone conducting or
administering statistical research specifically.
What Is Statistics? A Working Definition
Statistics is conventionally divided into two complementary halves. Descriptive
statistics summarizes and organizes a data set — means, medians, variances,
proportions, charts — without trying to generalize beyond it. Inferential
statistics goes further: it uses probability theory to quantify how confidently a
pattern observed in a sample can be generalized to the broader population it was drawn from,
and to distinguish a genuine effect from ordinary sampling noise. CASRAI covers this core
distinction in more depth in its guide to
descriptive vs.
inferential statistics.
Nearly every statistical method traces back to a small set of recurring questions:
- Design. How should data be collected so that the conclusions drawn from it
are actually valid — what should be measured, how should units be sampled or assigned to
conditions, and how large a sample is needed to detect an effect worth caring about? - Estimation. Given a sample, what is the best estimate of some unknown
quantity in the population (a mean, a proportion, a treatment effect), and how much uncertainty
surrounds that estimate? - Inference and testing. Is an observed pattern likely to reflect a real
effect, or could it plausibly have arisen from chance alone? CASRAI’s guides to
statistical significance and
statistical tests cover this machinery in
detail. - Modeling and prediction. What relationship, if any, connects two or more
variables, and how well can that relationship predict new, unobserved cases?
A statistician’s core skill is not running a specific test but choosing a design and a model
that actually match the data-generating process and the question being asked — misapplied
statistics (the wrong test, an unrepresentative sample, an uncontrolled confound) is one of the
most common sources of irreproducible research findings across every empirical field.
How Statistics Relates to Neighboring Disciplines
Statistics is unusual among academic disciplines in that almost every other empirical field
depends on it methodologically, while also developing its own statistical sub-specialty suited
to its own data:
- Mathematics. Mathematical statistics is built directly on probability
theory, measure theory, and linear algebra; many statistics departments originated inside, or
remain closely tied to, mathematics departments. - Computer science and data science. Statistical learning theory underlies
much of modern machine learning, and the applied field of
data science blends statistical inference with
computational methods and domain expertise. See CASRAI’s guide to
computer science. - Economics. Econometrics is, in large part, statistics applied to economic
data, sharing regression and causal-inference methods with applied statistics while developing
its own methodological tradition around observational economic data. See CASRAI’s guide to
economics. - Biology and medicine. Biostatistics applies statistical methods to
biological, clinical, and public-health data — clinical-trial design, survival analysis,
genetic-association studies — and has developed into its own recognized discipline with
distinct professional infrastructure. See CASRAI’s guide to
biostatistics. - Psychology and education. Psychometrics applies statistical modeling to
psychological measurement — test reliability and validity, latent-trait models. See
CASRAI’s guide to psychology. - Sociology, demography, and criminology. Quantitative social science leans
heavily on survey statistics, multilevel modeling, and demographic methods to study populations,
social structure, and crime patterns. See CASRAI’s guides to
sociology,
demography, and
criminology.
What distinguishes statistics from all of these applied fields is that it studies the
inferential machinery itself — the general theory of estimation, uncertainty, and
experimental design — rather than any one substantive domain’s data.
Major Subfields and Branches of Statistics
Statistics is a large discipline with many recognized subfields, ranging from purely
theoretical to intensely applied:
- Mathematical (theoretical) statistics — the probability-theoretic
foundations of estimation, hypothesis testing, and asymptotic theory that other subfields build
on. - Applied statistics — the practice of designing studies and analyzing
real data across a specific application domain, often embedded within another field’s own
methods literature. - Bayesian statistics — an inferential framework that treats unknown
quantities as having probability distributions updated by observed data, as opposed to the
frequentist framework built around long-run sampling properties; most modern statisticians work
comfortably in both traditions. - Biostatistics — statistical methods for biomedical, clinical, and
public-health research; see CASRAI’s dedicated guide to
biostatistics for more depth. - Survey statistics and sampling methodology — the design of probability
samples, weighting, and estimation for large population surveys, foundational to official
government statistics. - Time series analysis and forecasting — methods for data collected
sequentially over time, used across economics, climate science, and signal processing. - Spatial statistics — methods for data indexed by geographic location,
used in epidemiology, ecology, and geography. - Statistical genetics and genomics — methods for analyzing
high-dimensional genetic and genomic data, including genome-wide association studies. - Statistical computing and computational statistics — the algorithms
(Markov chain Monte Carlo, bootstrap resampling, optimization) that make modern statistical
inference computationally tractable; CASRAI’s guide to
bootstrapping in statistics covers one widely
used resampling technique. - Statistical learning and machine learning — the overlap between
statistics and computer science focused on prediction and pattern recognition from data,
increasingly central to both fields. - Official and government statistics — the production of national
economic, demographic, and social statistics by government statistical agencies. - Actuarial and financial statistics — statistical methods applied to
insurance risk, mortality, and financial markets. - Industrial and quality-control statistics — statistical process
control and reliability methods used in manufacturing and engineering.
Who Funds Statistics Research
Statistics research is funded primarily through federal science agencies, with additional
support from mission-specific government agencies and a smaller number of private
foundations:
- National Science Foundation (NSF). The Division of Mathematical
Sciences (DMS), part of NSF’s Directorate for Mathematical and Physical Sciences,
maintains a standing Statistics program that is the primary source of federal
funding for core statistical methodology and theory in the US. NSF’s Directorate for Social,
Behavioral and Economic Sciences also funds statistics through its
Methodology, Measurement and Statistics (MMS) program, housed in the Division
of Social and Economic Sciences, which specifically supports the development of statistical,
survey, and measurement methods used across the social and behavioral sciences. - National Institutes of Health (NIH). NIH does not maintain a single
dedicated statistics institute, but several NIH institutes fund the development of statistical
and biostatistical methodology as part of larger biomedical research grants — for example
methodology embedded in clinical-trial design, genetics, or epidemiology grants funded by the
relevant disease- or system-focused institute. See CASRAI’s guide to
biostatistics for more on how health-focused
statistical research is specifically funded. - Federal statistical agencies. The U.S. Census Bureau, the
Bureau of Labor Statistics (BLS), and the National Center for Health
Statistics (NCHS) are the core agencies of the US federal statistical system. They are
primarily producers and employers of applied statistical work rather than external grant-making
bodies, though the Census Bureau in particular also funds and conducts methodological research
(for example on disclosure-avoidance and differential-privacy methods) relevant to survey
statistics. - Private foundations. The Simons Foundation, through its
Mathematics and Physical Sciences division, is a major private funder of mathematical and
statistical research. The Alfred P. Sloan Foundation funds mathematics and data
science initiatives, including cross-institutional data-science research environments. As with
any funding landscape, program scope, budgets, and priorities shift — research offices
should confirm current guidelines directly on each funder’s own site before an application,
rather than relying on a general overview like this one.
Research Methods and Tools in Statistics
Statistical research and practice draws on a common toolkit, applied differently depending on
the question and the field:
- Study and experimental design — randomization, control groups,
blocking, and factorial designs intended to isolate a causal effect and support valid
inference, plus sample-size and power calculations to plan a study before data collection
begins. CASRAI’s guide to
statistical power
analysis and sample size covers this in more depth. - Sampling methodology — probability sampling designs (simple random,
stratified, cluster, multistage) used to select a representative subset of a population,
foundational to survey statistics and official statistics alike. - Descriptive and exploratory analysis — summarizing data through
measures of central tendency and dispersion and visualizing distributions before formal
modeling; see CASRAI’s guide to descriptive
statistics. - Hypothesis testing and estimation — the classical framework of
p-values, confidence intervals, and effect sizes, alongside degrees of freedom and other foundational
concepts that govern how these procedures are computed and interpreted. - Regression and generalized linear models — modeling how one or more
predictor variables relate to an outcome, extended to non-normal outcomes (binary, count) via
generalized linear models. - Resampling and simulation methods — bootstrap and permutation methods
that estimate uncertainty computationally rather than through closed-form formulas; see CASRAI’s
guide to bootstrapping in statistics. - Bayesian computation — Markov chain Monte Carlo and related
simulation methods used to fit Bayesian statistical models where no closed-form solution
exists. - Statistical software — R, Python (with libraries such as pandas,
statsmodels, and scikit-learn), SAS, Stata, and SPSS are the primary tools used across academic
and applied statistics; see CASRAI’s guide to SPSS for one widely used example.
Career and Training Pathways
Statistics has a fairly standardized academic training pipeline, with strong demand for
statisticians well beyond academia itself:
- Undergraduate. A bachelor’s degree in statistics or mathematics with a
statistics concentration typically covers probability theory, mathematical statistics,
regression, and at least one statistical-computing language, and serves as preparation for
graduate study or entry-level data-analyst and statistician roles. - Graduate training. A master’s degree (roughly one to two years) is a common
terminal degree for applied statistician roles in industry and government. A PhD in statistics
or biostatistics typically takes around four to six years: initial coursework in probability
theory and statistical inference culminating in qualifying exams, followed by original
dissertation research under a faculty advisor, usually developing new statistical methodology or
theory. - Professional societies. The American Statistical Association
(ASA), founded in 1839, is the principal professional society for statisticians in the
US, publishing journals including the Journal of the American Statistical Association
and offering a voluntary Accredited Professional Statistician (PStat®)
credential. The Institute of Mathematical Statistics (IMS), founded in 1935,
is the leading international society for theoretical and mathematical statistics, publishing the
Annals of Statistics. The Royal Statistical Society (RSS) plays the
equivalent role in the UK. CASRAI’s guide to biostatistics also covers the
International Biometric Society, the professional
home for statisticians working in the biosciences specifically. - Career destinations. Statisticians and biostatisticians work across
academia, federal statistical agencies (the Census Bureau, BLS, and NCHS among them),
pharmaceutical and biotechnology companies, technology companies (as data scientists and machine
learning engineers), insurance and finance (as actuaries and quantitative analysts), and
government policy agencies — reflecting how thoroughly statistical training transfers
across sectors.
Frequently Asked Questions
What is the difference between statistics and data science?
Statistics is the older, foundational discipline that develops the theory of inference,
estimation, and uncertainty; data science is a
newer, more applied field that combines statistical methods with computer science and
domain-specific expertise to extract insight from data, often at larger scale and with a
heavier emphasis on computation and software engineering than traditional statistics
training.
What is the difference between statistics and biostatistics?
Biostatistics is the branch of statistics specifically focused on biological, clinical, and
public-health data — clinical-trial design, survival analysis, epidemiological methods.
General statistics covers the same inferential theory applied across every domain, not just
health and biology; see CASRAI’s dedicated guide to
biostatistics for more detail.
Do you need a PhD to work as a statistician?
No. Many applied statistician and data-analyst roles in industry and government are held by
people with a bachelor’s or master’s degree in statistics; a PhD is generally expected for
academic research positions and for roles developing genuinely new statistical methodology, but
is not a prerequisite for applied statistical practice.
What math is required for statistics?
Undergraduate statistics typically requires calculus and linear algebra; graduate-level and
PhD statistics requires substantially more — probability theory, mathematical analysis,
and measure theory for the more theoretical subfields — since most statistical inference
is built on that mathematical foundation.
Is statistics part of mathematics?
Statistics grew out of mathematics and remains grounded in probability theory, but it is
widely treated as a distinct discipline in its own right because so much of its content —
study design, applied data analysis, domain-specific methodology — is not purely
mathematical. Many universities have separate statistics and mathematics departments for exactly
this reason.
Related CASRAI Resources
This guide is part of CASRAI’s Branches of Science
series, which maps the major academic and scientific disciplines along with the
research-administration context — funding, methods, and career pathways — that a
general encyclopedia entry typically omits. Related discipline guides include
biostatistics,
data science,
economics,
criminology, and
demography, plus CASRAI’s broader research-methods
guides on descriptive statistics,
statistical significance, and
bootstrapping for deeper treatment of specific
statistical methods.








