Skip to main content
v2026.11,858 entries · CC-BY 4.0

What Is Machine Learning? Research Methods, Pitfalls, and Funding

What machine learning is, its subfields, how to evaluate models without data leakage, who funds the research (NSF), and the key conferences and journals.

Written and maintained by CASRAI Editorial Board

Last updated

Machine learning (ML) is the field concerned with algorithms that improve their performance on a task by learning from data rather than by following rules written out explicitly by a programmer. The most widely quoted formal definition comes from Tom Mitchell’s 1997 textbook Machine Learning: a program is said to learn from experience E with respect to a class of tasks T and a performance measure P if its performance at tasks in T, as measured by P, improves with experience E. The definition is useful because it forces three questions that every ML study has to answer: what is the task, what counts as good performance, and what data counts as experience. Most of the methodological problems discussed later on this page come from answering one of those three questions carelessly.

Machine learning is a subfield of artificial intelligence and draws on statistics, mathematics (linear algebra, probability, and optimization in particular), and computer science. It is also the technical core of data science. Increasingly it is a research method in its own right, used by biologists, chemists, social scientists, clinicians, and physicists to build predictive models of their own data, which is why this page spends as much time on how ML results are evaluated and reproduced as on what the algorithms are.

What Machine Learning Studies

ML research asks a small set of recurring questions regardless of application: how can a model be fit to data efficiently, how well will it perform on data it has not seen (generalization), how much data is needed to learn a given task, how can a model’s behavior be understood or explained, and how can a learned system be made robust when conditions change. The field’s theoretical wing (statistical learning theory, optimization theory) asks these questions mathematically; its empirical wing answers them by running experiments on benchmark datasets and real-world problems.

Three ideas recur across nearly every ML method:

  • A model with adjustable parameters. A model is a function whose behavior is set by parameters, from a handful in a linear regression to billions in a large neural network.
  • A loss function. A numerical measure of how wrong the model’s outputs are on the training data, which the learning algorithm tries to reduce, usually with a form of gradient-based optimization.
  • Generalization. The goal is never low error on the data used for fitting; it is low error on new data drawn from the situation the model will actually be used in. The gap between the two is overfitting, and managing it is the central practical concern of the field.

Major Subfields and Learning Paradigms

  • Supervised learning — learning to predict a known outcome (a label or a number) from labeled examples. Classification (is this scan malignant?) and regression (what will this value be?) are the two basic forms, and most applied ML in research falls here.
  • Unsupervised learning — finding structure in data without labels: clustering, dimensionality reduction, density estimation, and anomaly detection.
  • Semi-supervised and self-supervised learning — using a large amount of unlabeled data, plus a smaller amount of labeled data or labels derived from the data itself, to learn useful representations. Self-supervised pretraining underlies most current large language models.
  • Reinforcement learning — learning a strategy for sequential decisions from reward signals gathered by interacting with an environment. Reinforcement learning from human feedback is covered in CASRAI’s dictionary entry on RLHF.
  • Deep learning — ML built on multi-layer artificial neural networks, which learn their own intermediate representations of raw inputs such as pixels, audio, or text. It dominates computer vision, speech, and natural language processing.
  • Probabilistic, Bayesian, and kernel methods — approaches that represent uncertainty explicitly, or that use similarity functions to handle nonlinear problems without a deep network. They remain important where data are scarce or calibrated uncertainty matters.
  • Causal and interpretable ML — work on distinguishing prediction from causal effect, and on explaining why a model produced a given output; see CASRAI’s guide to explainable AI.
  • Trustworthy ML — fairness, privacy, robustness, and safety of learned systems. The governance side of this work is covered in CASRAI’s guide to AI governance.

Application areas such as computer vision, natural language processing, robotics (see what is robotics), and computational biology (see what is bioinformatics) are usually treated as neighbors of ML rather than subfields, since each has its own data, benchmarks, and communities.

A Brief History

The phrase “machine learning” is usually credited to Arthur Samuel, whose 1959 paper described a checkers-playing program that improved by playing against itself. Frank Rosenblatt’s perceptron, from the late 1950s, was an early trainable neural model. Through the following decades the field alternated between enthusiasm and disappointment, with symbolic and rule-based AI dominating for long stretches. The backpropagation algorithm, popularized by Rumelhart, Hinton, and Williams in a 1986 Nature paper, made training multi-layer networks practical, and the 1990s brought support vector machines, ensemble methods, and a more statistical orientation to the field.

The modern deep-learning era is commonly dated to 2012, when the AlexNet convolutional network won the ImageNet image-classification competition by a wide margin, helped by large labeled datasets and GPU computing. The transformer architecture, introduced in the 2017 paper “Attention Is All You Need,” later became the basis of today’s large language models. Throughout this history, progress has depended as much on data, compute, and benchmarks as on new algorithms, which is one reason evaluation practice (next section) matters so much.

ML as a Research Method: Evaluation and Reproducibility Pitfalls

When a paper reports that an ML model predicts something with high accuracy, the claim is only as good as the evaluation behind it. Evaluation in supervised learning rests on one rule: the data used to assess the model must be independent of the data used to build it. Standard practice therefore splits data into training, validation (used for choosing hyperparameters and comparing model variants), and test sets, the last touched once, at the end. Cross-validation repeats the split several times to get a more stable estimate when data are limited.

Data Leakage

Data leakage occurs when information that would not be available at prediction time finds its way into model training or model selection, so that performance on the test set is inflated. CASRAI has a dictionary entry for data leakage (training). Typical forms include:

  • Preprocessing before splitting. Normalizing, imputing, or selecting features using the whole dataset and only then splitting lets test-set information shape the training pipeline.
  • Duplicates and non-independent samples. Multiple records from the same patient, subject, or document landing on both sides of the split, so the model recognizes the individual rather than learning the general pattern.
  • Temporal leakage. Using future data to predict the past, for example training on later observations and testing on earlier ones.
  • Illegitimate features. Including a variable that is itself a consequence of, or a proxy for, the outcome being predicted.
  • Test set not representative of the target use. Evaluating on data that differ from deployment conditions, so that the estimate does not apply to the real task.

A study by Sayash Kapoor and Arvind Narayanan, published in the journal Patterns in 2023 under the title “Leakage and the reproducibility crisis in machine-learning-based science,” surveyed papers across scientific fields and found evidence of data leakage in 294 papers across 17 fields. It proposed a taxonomy of eight types of leakage and recommended “model info sheets” as a way for authors to document why leakage has been ruled out. In a case study on civil war prediction, the apparent superiority of complex ML models over older statistical models disappeared once the leakage was corrected. The practical lesson is that a surprisingly high score deserves suspicion before celebration.

Other Common Pitfalls

  • Weak baselines. Reporting that a new model beats an older one without tuning the older one with comparable effort, or without including a simple baseline such as a regularized linear model.
  • Test-set reuse. Repeatedly consulting the test set while developing a model turns it into a second validation set, and the final reported number is then optimistic. Community benchmarks suffer from this at scale, since many papers tune against the same test data.
  • Uncertainty ignored. Reporting a single run’s score without variation across random seeds, data splits, or confidence intervals, so that differences of a point or two cannot be told apart from noise.
  • Inappropriate metrics. Accuracy on a highly imbalanced dataset can look excellent while the model finds none of the rare cases that matter; the metric has to match the decision the model supports.
  • Distribution shift. A model validated at one hospital, on one instrument, or in one population can fail elsewhere; external validation on independent data is the stronger test.
  • Computational irreproducibility. Results that depend on unreported random seeds, library versions, hardware, or undisclosed code and trained weights cannot be independently checked.

Reproducibility Practices

The ML community has responded with checklists and norms. Major venues now ask authors to complete reproducibility checklists at submission; Joelle Pineau and colleagues described the experience of introducing such a checklist at NeurIPS in a 2021 paper in the Journal of Machine Learning Research. Good practice includes releasing code and configuration, fixing and reporting random seeds, versioning data and software environments, documenting datasets and models, and reporting results over several runs. CASRAI’s dictionary covers the concept of a reproducible AI experiment, and the same research data management practices apply as in any computational field: a data management plan, the FAIR data principles, and attention to training data provenance. For the wider context see the entry on the reproducibility crisis.

In health research, reporting guidelines exist for prediction-model studies, including TRIPOD+AI and the others compared in CASRAI’s guide to AI reporting guidelines, and the MINIMAR checklist for clinical ML models. STARD and TRIPOD cover diagnostic-accuracy and prediction-model reporting more generally. Pre-specifying the evaluation plan before looking at test results, much as pre-registration does in other fields, is a further safeguard.

Who Funds Machine Learning Research

In the United States the National Science Foundation (NSF) is the main federal funder of academic ML research, principally through the Directorate for Computer and Information Science and Engineering (CISE; see NSF CISE). CISE’s Division of Information and Intelligent Systems (IIS) supports AI and ML broadly, and statistical foundations are supported through NSF’s Division of Mathematical Sciences. NSF also funds the National AI Research Institutes, a multi-year program run under successive solicitations (NSF 20-503 was the first, and NSF 23-610 is the most recent version we could confirm); see NSF AI Institutes. The National AI Research Resource (NAIRR) aims to widen researcher access to compute and data.

Federal programs are reorganized and renamed often, and solicitations expire, so confirm the current program and deadline on NSF’s own site before building a proposal around any of the names above. Other U.S. agencies, including NIH and the Department of Energy, fund ML as applied to their own missions, and private foundations and industry laboratories are unusually large participants in ML specifically; a considerable share of frontier ML research happens outside universities. For the policy side, see CASRAI’s dictionary entries on NSF AI Policy and NIH AI Policy.

Research administrators supporting ML grants should be aware of the field’s particular cost structures: compute and cloud credits, data acquisition and licensing, and data-sharing agreements. CASRAI’s research administration hub and grants management hub cover the pre- and post-award work that surrounds these awards, and the guide on the NIST AI Risk Management Framework for research institutions covers one governance framework institutions increasingly apply.

Conferences, Journals, and Societies

Unlike many sciences, ML research is published primarily at peer-reviewed conferences, where full papers are reviewed and archived, with journals playing a complementary role.

  • NeurIPS (Conference on Neural Information Processing Systems) began in 1987 under founding general chair Ed Posner and is among the largest and most influential venues in the field.
  • ICML (International Conference on Machine Learning) traces to a 1980 workshop at Carnegie Mellon University and has used the conference name and numbering since 1988.
  • ICLR (International Conference on Learning Representations) is, with NeurIPS and ICML, usually counted among the field’s three leading general venues, with a focus on representation learning and deep learning.
  • The Journal of Machine Learning Research (JMLR) published its first papers in October 2000 as an open-access journal, after editorial board members of an older subscription journal, Machine Learning, resigned in protest at paywalled archives.
  • Other well-known venues include COLT (learning theory), AISTATS, and KDD (data mining, run by ACM SIGKDD), and the journals Machine Learning and IEEE Transactions on Pattern Analysis and Machine Intelligence. Application-specific ML work often appears in the journals of its home discipline.

Professional societies include the Association for Computing Machinery (ACM), the IEEE and its Computer Society, and the Association for the Advancement of Artificial Intelligence (AAAI). Because conference papers carry more weight in this field than in most, and preprints on arXiv are standard practice, authors should check how a funder or institution counts conference papers and preprints in evaluation.

Training Paths and Careers

People enter ML research through computer science, statistics, mathematics, electrical engineering, or a quantitative domain science. Many universities now offer dedicated degrees or concentrations in machine learning or AI at master’s and PhD level. A research PhD typically combines coursework in probability, linear algebra, optimization, and statistical learning with an original dissertation, and publication at the field’s main conferences is a normal expectation during the degree. Careers span academia, industry research laboratories, applied ML engineering, national laboratories, and domain scientists who use ML as a tool.

Frequently Asked Questions

What is machine learning in one sentence?
Machine learning is the study of algorithms that improve at a task by learning patterns from data instead of following hand-written rules.

What is the difference between machine learning and artificial intelligence?
AI is the broader goal of building systems that perform tasks associated with intelligence, and includes approaches such as logic and search that involve no learning. Machine learning is the subfield that does this by learning from data, and it is the approach behind most current AI systems. See CASRAI’s guide to artificial intelligence.

What is the difference between machine learning and deep learning?
Deep learning is a subset of machine learning that uses multi-layer neural networks. Other ML methods, such as decision-tree ensembles, linear models, and kernel methods, are not deep learning and are often competitive on small or tabular datasets.

What is the difference between machine learning and statistics?
They share much mathematics. Classical statistics has traditionally emphasized inference about parameters and quantified uncertainty; ML has emphasized prediction accuracy on new data, often with complex models at scale. The boundary is blurry and the fields increasingly borrow from each other. See what is statistics.

What is data leakage in machine learning?
It is the contamination of model training or selection with information that would not be available when the model is used, which inflates test performance. Examples are preprocessing before the train/test split, the same subject appearing in both sets, and using future information to predict the past.

Why do train, validation, and test sets matter?
The training set fits the model, the validation set is used to make choices such as hyperparameters, and the test set gives a final estimate of performance on new data. Using the test set for any decision makes that estimate optimistic.

Where is machine learning research published?
Chiefly at conferences such as NeurIPS, ICML, and ICLR, and in journals such as the Journal of Machine Learning Research, with domain applications published in discipline-specific journals.

Who funds machine learning research in the United States?
Mainly NSF (especially CISE), alongside mission agencies such as NIH and DOE, private foundations, and industry laboratories. Check each funder’s current programs directly.

Where Machine Learning Fits Among the Sciences

For the wider map of disciplines, see CASRAI’s overview of the branches of science. Closely related discipline guides in the same series include artificial intelligence, data science, statistics, computer science, and cognitive science.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about What Is Machine Learning? Research Methods, Pitfalls, and Funding

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.