Written and maintained by CASRAI Editorial Board
Last updated
Data science is the interdisciplinary field that uses statistical methods, algorithms, computational systems, and domain expertise to extract knowledge, patterns, and actionable insight from data — structured and unstructured, at any scale. It sits at the intersection of statistics, computer science, and a specific application domain: a data scientist working on genomics, financial markets, or public-health surveillance draws on much of the same core statistical-learning and computational toolkit, applied to very different data and questions. That three-way blend — computational and programming skill, mathematical and statistical knowledge, and substantive domain expertise, sometimes illustrated as three overlapping circles — is what most working definitions of the field converge on, and it is also part of why data science organizes and funds itself differently than a traditional single-department discipline.
What Data Science Actually Studies
Data science asks a recurring set of questions regardless of the domain it is applied to: What does this data actually contain, and how was it collected and structured? What patterns or relationships in it are real signal rather than noise? Can a model built on past data predict or generalize to new cases reliably? And, where a decision or intervention is at stake, does an observed association reflect an actual causal effect, or just correlation?
- Statistical inference and uncertainty quantification — estimating quantities of interest from data and characterizing how confident that estimate should be, the same core statistical logic classical statistics is built on, applied at data scales and structures that classical methods were not originally designed for.
- Machine learning — algorithms that learn patterns directly from data rather than following explicitly programmed rules, spanning supervised learning (predicting a known outcome from labeled examples), unsupervised learning (finding structure in unlabeled data, such as clustering), and reinforcement learning (learning a strategy through trial-and-error feedback).
- Data engineering — the pipelines, infrastructure, and systems that collect, clean, validate, store, and move data reliably at scale, without which the analysis and modeling layers above have nothing trustworthy to work from.
- Data mining and knowledge discovery — systematically searching large datasets for patterns, associations, or anomalies that were not specified in advance.
- Data visualization and communication — representing data and findings so that patterns are actually interpretable and results are usable by a decision-maker who was not part of the analysis.
- Experimentation and causal inference — designing controlled experiments (such as A/B tests) and applying quasi-experimental statistical methods to support causal claims from observational data, where a controlled experiment is not possible.
A Brief History of Data Science
The term itself predates its current meaning by decades. Computer scientist Peter Naur used “data science” in his 1974 book Concise Survey of Computer Methods, largely as a synonym for computer science as it was then understood. John Tukey’s 1962 essay “The Future of Data Analysis” is widely cited as an important conceptual precursor: Tukey argued that data analysis deserved recognition as its own empirical science, not merely a technical subset of mathematical statistics. Statistician William S. Cleveland’s 2001 paper “Data Science: An Action Plan for Expanding the Technical Areas of the Field of Statistics” proposed a formal expansion of the statistics discipline into what he explicitly called data science, complete with a proposed curriculum.
The job title “data scientist” is usually credited to DJ Patil and Jeff Hammerbacher, who independently began using the term around 2008 while building data teams at LinkedIn and Facebook, respectively. Harvard Business Review’s 2012 article “Data Scientist: The Sexiest Job of the 21st Century” (Davenport and Patil) brought the label to wide public and business attention. Through the 2010s, cheaper storage and computation, growth in the volume and variety of available data, and advances in machine learning — particularly the deep-learning resurgence that accelerated sharply after 2012 — drove data science’s institutionalization as a named academic field with its own degree programs, university departments, and, increasingly, dedicated funding lines distinct from its parent disciplines of statistics and computer science.
How Data Science Relates to Neighboring Disciplines
Data science’s closest and most contested boundary is with statistics: both fields share the same inferential and statistical-learning core, and much of what is taught in a data science program is, formally, statistics. The distinguishing emphasis in practice tends to be data science’s greater weight on computation at scale, engineering pipelines for messy or unstructured data (text, images, sensor streams, logs), and building systems that operationalize a model in production, rather than statistics’ traditional emphasis on formal inferential theory and experimental design.
Data science’s other close boundary is with computer science, which supplies much of the field’s algorithmic and systems foundation — machine learning itself grew up largely inside computer science departments, and CASRAI’s companion computer science guide notes the same overlap from the other side. Mathematics underlies both fields’ methods directly, through linear algebra, optimization, and probability theory. Beyond these two, data science functions as an applied methodology layered onto other disciplines’ own data and questions: bioinformatics applies data-science methods to biological and genomic data; epidemiology and public-health surveillance increasingly depend on the same large-scale statistical-learning toolkit; economics‘s econometrics tradition and political science‘s computational social science subfield both draw on data science methods applied to economic and social data; and cognitive science and artificial intelligence research overlap substantially with data science’s machine-learning core.
Major Sub-disciplines and Subfields of Data Science
Most data science programs and research groups organize around some version of the following areas, though naming and boundaries vary by institution:
- Machine learning and statistical learning — building and evaluating predictive and pattern-recognition models, including the deep-learning methods (built on artificial neural networks) that currently dominate the field’s fastest-growing applications.
- Data mining and pattern discovery — extracting previously unknown, potentially useful patterns and relationships from large datasets.
- Big data and distributed computing systems — storing, processing, and analyzing datasets too large or fast-moving to handle on a single machine.
- Data engineering and MLOps — building and maintaining the pipelines and infrastructure that move data reliably from source to model to production, and that keep a deployed model working correctly over time.
- Computational statistics — the statistical methods and computational techniques (resampling, Bayesian computation, high-dimensional inference) that make modern statistical learning practical at scale.
- Natural language processing and computer vision — applied subfields extracting structure and meaning from text and image or video data specifically, sitting squarely across the data science/computer science boundary.
- Data visualization and visual analytics — the design and study of visual representations that make complex data interpretable.
- Data ethics and responsible AI — the study of fairness, bias, privacy, and accountability in data-driven systems and algorithmic decision-making, a fast-growing subfield as automated decisions affect more of daily life.
- Applied domain data science — health and biomedical data science, computational social science, business and marketing analytics, and geospatial data science are among the most active application areas, each layering data science methods onto an existing domain’s own questions and data.
Who Funds Data Science Research
In the United States, the National Science Foundation (NSF) is the primary federal funder of academic data science research, spread across several directorates rather than concentrated in one: the Directorate for Computer and Information Science and Engineering (CISE) funds the field’s algorithmic and computational foundations (see CASRAI’s NSF CISE dictionary entry), the Directorate for Mathematical and Physical Sciences (MPS)‘s Division of Mathematical Sciences funds the statistical and mathematical foundations, and the Directorate for Social, Behavioral and Economic Sciences (SBE) funds social and behavioral data science specifically. NSF has also run a cross-directorate “Harnessing the Data Revolution” (HDR) Big Idea, which funded a national network of NSF HDR Institutes for data-science research beginning in 2019; as with CISE’s own internal structure, NSF periodically reorganizes its cross-cutting Big Idea initiatives, so researchers should verify the current program structure directly with NSF rather than assume the exact institute list is unchanged.
On the biomedical side, the National Institutes of Health (NIH)‘s Office of Data Science Strategy (ODSS), established in September 2022, coordinates data science strategy and shared infrastructure — common capabilities, metadata, and metrics across NIH’s generalist and domain-specific data repositories — building on the earlier Big Data to Knowledge (BD2K) initiative that NIH ran from 2012 to 2017. Individual NIH institutes also fund data science methods applied within their own biomedical focus areas, with the National Library of Medicine playing a particularly direct role in biomedical informatics and health data science. The Department of Energy’s Office of Advanced Scientific Computing Research (ASCR) funds the computational and data-intensive science infrastructure, much of it based at DOE national laboratories, that large-scale data science research depends on.
Among private foundations, the Alfred P. Sloan Foundation and the Gordon and Betty Moore Foundation jointly funded the Moore-Sloan Data Science Environments initiative, which established dedicated data science institutes at the University of Washington, New York University, and the University of California, Berkeley beginning in 2013 — a genuinely notable early philanthropic commitment to data science as an academic field in its own right, distinct from either statistics or computer science departments. The Simons Foundation, primarily known for funding theoretical computer science and mathematics, also supports large-scale computational research relevant to data science through its Flatiron Institute. Newer entrants such as Schmidt Sciences fund large-scale AI and computational research initiatives that overlap substantially with data science. As with any funder, researchers should verify current program details and eligibility directly rather than rely on a static list. As in computer science, private-sector research labs and sponsored-research funding also play an unusually large role specifically in data science, given the field’s direct commercial applications.
Typical Research Methods, Tools, and Practices
Data science research is built on a common technical stack: statistical modeling and analysis (heavily using the R and Python programming languages, including Python’s scientific-computing and statistics libraries), machine learning frameworks for building and training predictive models, and distributed computing frameworks such as Apache Hadoop and Apache Spark for datasets too large to process on a single machine. Data is typically stored and queried through relational (SQL) or non-relational (NoSQL) database systems, and analysis is usually conducted in reproducible, notebook-based computational environments (such as Jupyter notebooks) tracked under version control (such as Git), so that an analysis can be re-run and checked by someone other than its original author.
Because data science research itself generates and depends on large, often sensitive datasets, the same research-data-management practices CASRAI documents elsewhere apply directly to data science research groups, not only to research-data-management specialists: a formal data management plan for how a project’s data will be collected, stored, and shared; adherence to the FAIR data principles (Findable, Accessible, Interoperable, Reusable) when datasets are deposited for reuse; and, for datasets deposited in a public repository, certification such as CoreTrustSeal as a trust signal for that repository. Reproducibility is a live methodological concern specific to computational and machine-learning research in particular — not just the underlying data, but the exact code, trained model parameters, and software environment used to produce a result increasingly need to be preserved and shared for a computational result to be independently checked, a domain-specific instance of the broader reproducibility concerns CASRAI covers across the research-integrity literature.
Career and Training Pathways
Many universities now offer a data science degree directly — bachelor’s, master’s, and, increasingly, PhD programs — alongside the older and still common path of entering the field through a Statistics, Computer Science, or quantitative domain-science degree with a data science concentration or track. A research-track PhD in data science generally follows the same broad structure as other quantitative sciences: coursework in statistics, machine learning, and programming in the first one to two years, a qualifying examination, and then original dissertation research conducted under a faculty advisor, culminating in a written dissertation and defense.
The field’s relevant professional societies include the American Statistical Association (ASA), whose Statistical Learning and Data Science Section organizes work specifically at the statistics/data-science boundary; the Association for Computing Machinery (ACM), whose Special Interest Group on Knowledge Discovery and Data Mining (SIGKDD) organizes one of the field’s leading applied research conferences; the IEEE Computer Society, which organizes IEEE Big Data and related conferences; and INFORMS (the Institute for Operations Research and the Management Sciences), whose Certified Analytics Professional (CAP) credential is a well-known voluntary professional certification — data science has no state licensure requirement comparable to fields such as engineering or medicine. Career paths span academic research, industry data science and machine learning engineering roles, national laboratories, and government statistical and data agencies; as in computer science, a large share of data science PhD graduates move directly into industry research roles rather than a traditional academic postdoctoral position, a pattern the two fields share and that distinguishes them from many of the other sciences covered in this series.
Frequently Asked Questions
What is data science, in one sentence?
Data science is the interdisciplinary field that combines statistical methods, computer science and machine learning, and domain expertise to extract knowledge and actionable insight from data at any scale.
What is the difference between data science and statistics?
The two share the same inferential and statistical-learning core; data science places relatively more emphasis on computation at scale, engineering pipelines for messy or unstructured data, and deploying models in production, while classical statistics places relatively more emphasis on formal inferential theory and experimental design. Many data science programs are, formally, an applied extension of statistics.
What is the difference between data science and computer science?
Data science is generally treated as an applied, interdisciplinary field that draws on computer science (algorithms, databases, machine learning), statistics, and domain expertise to extract insight from data; computer science is the broader discipline data science draws several of its core computational methods from, alongside data science’s equally significant statistical and domain-specific components — see CASRAI’s computer science guide for the same distinction from the other side.
Is data science a STEM field, and does it require a PhD?
Yes, it is a STEM field, but a PhD is not required for most data science careers — many practicing data scientists hold a bachelor’s or master’s degree, with a PhD more common specifically among those pursuing academic or industry research-scientist roles rather than applied data-science roles.
How is data science research typically funded in the United States?
Mainly through the National Science Foundation, spread across its CISE, MPS, and SBE directorates plus its cross-directorate HDR Big Idea; the NIH’s Office of Data Science Strategy and individual biomedical institutes for health-focused work; the DOE’s Office of Advanced Scientific Computing Research for large-scale computational infrastructure; and private foundations including the Alfred P. Sloan Foundation, the Gordon and Betty Moore Foundation, and the Simons Foundation. See the funding section above for how these sources divide the field’s territory.
What tools and programming languages do data scientists use?
Most commonly Python and R for statistical analysis and modeling, machine learning frameworks for building predictive models, Apache Hadoop and Apache Spark for datasets too large for a single machine, SQL and NoSQL databases for storage and querying, and Jupyter notebooks with Git version control for reproducible analysis.
Where Data Science Fits Among the Sciences
For a broader map of how data science relates to the full set of major scientific disciplines — from physics and mathematics through to biology and the social sciences — see CASRAI’s overview guide to the branches of science, which this page is part of a companion series alongside. Related guides in that series include what is computer science, what is cognitive science, and what is epidemiology, each of which shares a genuine methodological or infrastructural border with data science; the concurrent what is public health and what is biomedical engineering guides in this same series (at what is public health and what is biomedical engineering) also draw directly on data science methods for health-data-intensive research once published.








