Skip to main content
v2026.11,858 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

Training data composition

The actual makeup of a training corpus used to build an AI model -- the mix of languages, domains, source types (licensed, scraped, user-generated, synthetic), and demographic or geographic representation it contains -- as distinct from where that data came from (see training data provenance).

ByCASRAI Editorial Board
· Last updated 5 Sept 2026
Share this

Ask CASRAI · free to try

Ask about Training data composition

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    A model card reporting the approximate proportion of code, academic text, and web-crawled text in a training corpus

  • Is an instance

    An audit finding a language model's training data is heavily skewed toward English and a small number of high-resource languages, with corresponding performance gaps in others

Counter-examples

Looks similar, but isn't

  • Not an instance

    A record of which organisations supplied a training dataset and under what licence is training data provenance, not composition

  • Not an instance

    A model's output-level accuracy disparity across groups is ai-bias -- a downstream effect that composition can contribute to but is not itself a description of composition

Editorial commentary

Training data composition describes the content mix of a model’s training corpus: what proportion is code versus prose, which languages and domains are represented and in what balance, how much is licensed or curated content versus unfiltered web scrape, and increasingly, how much is itself AI-generated (synthetic data) rather than human-produced.

Why composition matters downstream

Composition is one of the most direct upstream causes of a model’s behaviour and limitations. A corpus skewed toward high-resource languages, certain geographic regions, or particular demographic groups tends to produce a model that performs less reliably, or reproduces skewed assumptions, for underrepresented groups and use cases — one of the mechanisms behind AI bias. Composition also interacts directly with contamination risk: the more a training corpus draws indiscriminately from the open web, the higher the chance that public benchmark datasets or evaluation sets were incidentally swept in, inflating reported performance — see data leakage (training) and the documented contamination findings on benchmarks such as MMLU.

How this differs from related AI-band terms

  • vs. training data provenance: composition is a description of content and mix; provenance is a record of origin, licensing basis, and chain of custody. A developer could disclose full provenance (every source named) while still declining to characterise the resulting composition, or vice versa — the two are complementary, not interchangeable.
  • vs. AI fairness: composition is a factual, descriptive property of the input data; fairness is a normative judgment about whether resulting outcomes are acceptable, evaluated using contested, mutually incompatible formal criteria.

Documentation practice

Composition is typically reported, where it is reported at all, through structured documentation frameworks — datasheets for datasets, model cards, and data statements for NLP — rather than through a single standardised regulatory disclosure, though the EU AI Act’s Article 53 training-data summary requirement for general-purpose AI models pushes providers toward more systematic reporting of this kind.

Also known as

data mix · training mixture

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Training data composition"
      vocab-term-identifier="https://casrai.org/dictionary/term/training-data-composition" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/training-data-composition",
  "name": "Training data composition",
  "identifier": "https://casrai.org/dictionary/term/training-data-composition",
  "description": "The actual makeup of a training corpus used to build an AI model -- the mix of languages, domains, source types (licensed, scraped, user-generated, synthetic), and demographic or geographic representation it contains -- as distinct from where that data came from (see training data provenance).",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/training-data-composition",
  "alternateName": [
    "data mix",
    "training mixture"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:51",
  "dateModified": "2026-09-05T14:24:43",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

Research-admin question? Get an answer that links its sources.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.