Skip to main content
v2026.11,858 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

Data statement (NLP)

A standardised description of an NLP dataset covering curation rationale, language variety, speaker and annotator demographics, speech situation, text characteristics, and recording quality.

ByCASRAI Editorial Board
· Last updated 5 Sept 2026
Share this

Ask CASRAI · free to try

Ask about Data statement (NLP)

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    A coreference dataset published with a data statement specifying US Mainstream English (white, college-educated, ~30-50yo annotators).

  • Is an instance

    A speech-recognition corpus's data statement specifying broadcast news from 1995-2005.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A token count and vocabulary size summary.

  • Not an instance

    A model card (describes the model, not the data).

Editorial commentary

A data statement is a standardised description of a natural-language-processing dataset, proposed by Bender and Friedman in “Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science” (Transactions of the ACL, 2018). It covers curation rationale, language variety, speaker and annotator demographics, the speech situation the text was produced in, text characteristics, and recording quality — fields chosen specifically because their absence had repeatedly let biases (most commonly, datasets skewed toward a narrow demographic or a single English variety) go unnoticed until a model trained on them failed for other populations.

How this differs from a datasheet

A datasheet for datasets (Gebru et al., 2018/2021) is the general-purpose ML-dataset documentation template, covering motivation, composition, collection process, and uses across any data modality. A data statement is the NLP-specific counterpart, developed independently and slightly earlier, with fields tuned specifically to the sociolinguistic properties — speaker demographics, language variety, speech situation — that matter for text data but have no equivalent in, say, an image dataset. In practice a project working with text data may complete a data statement instead of, or alongside, a general datasheet.

Why it matters

Reviewers assessing whether an NLP model or dataset is fit for a particular population should be able to check the data statement’s language-variety and demographic fields directly, rather than inferring coverage from the model’s stated aggregate performance alone.

References

Also known as

NLP data statement · Bender-Friedman data statement

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Data statement (NLP)"
      vocab-term-identifier="https://casrai.org/dictionary/term/data-statement-nlp" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/data-statement-nlp",
  "name": "Data statement (NLP)",
  "identifier": "https://casrai.org/dictionary/term/data-statement-nlp",
  "description": "A standardised description of an NLP dataset covering curation rationale, language variety, speaker and annotator demographics, speech situation, text characteristics, and recording quality.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/data-statement-nlp",
  "alternateName": [
    "NLP data statement",
    "Bender-Friedman data statement"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:50",
  "dateModified": "2026-09-05T17:31:40",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

Research-admin question? Get an answer that links its sources.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.