Skip to main content
v2026.11,858 entries · CC-BY 4.0

Model Cards, System Cards, and the EU Model Documentation Form

Model cards, system cards and the EU Model Documentation Form are three different instruments with three different levels of legal force. What each actually requires, which one is mandatory, and where research institutions pick up obligations of their own.

Written and maintained by CASRAI Editorial Board

Last updated

Three documents in frontier AI governance have names that sound almost interchangeable, and they are not. A model card is a research-community convention from 2018. A system card is what frontier labs actually publish when they ship. The Model Documentation Form is a European Commission instrument that most people will never see, because it is not published at all. Only one of the three is close to mandatory, and it is not the one that gets cited most. This guide separates them, and notes where CASRAI’s own NIKOLAI dictionary proposes a field — redaction, in track N8 — for the thing all three leave unrecorded.

The short version

Instrument Origin Who sees it Legal force
Model card Mitchell et al., academic paper, 2018 Public, if published None. A convention.
System card Frontier lab practice (OpenAI, Anthropic) Public None directly, but named in California statute as an acceptable container
Model Documentation Form EU GPAI Code of Practice, Transparency chapter, 10 July 2025 AI Office, national authorities, downstream providers — on request Voluntary code, evidencing a binding AI Act duty
Public summary of training content EU AI Office template, adopted 24 July 2025 Public Mandatory under AI Act Article 53(1)(d)

The model card: a 2018 research proposal, never codified

The model card originates in Model Cards for Model Reporting by Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji and Timnit Gebru, posted to arXiv on 5 October 2018 (arXiv:1810.03993) and presented at the ACM Conference on Fairness, Accountability, and Transparency (FAT* ’19) in Atlanta in January 2019.

Its central argument is disaggregation: a single headline accuracy number hides the fact that a model may perform very differently across demographic and phenotypic groups, and across the intersections of those groups. The paper’s worked examples are a face-smile detector and a toxic-comment classifier, both benchmarked by group rather than in aggregate.

The paper proposes nine sections:

  1. Model Details — developer, version date, model type, architecture.
  2. Intended Use — primary use cases, intended users, out-of-scope applications.
  3. Factors — the groups, instrumentation and environments across which performance is expected to vary.
  4. Metrics — the measures chosen, and why.
  5. Evaluation Data — what the model was assessed on.
  6. Training Data — what it was trained on.
  7. Quantitative Analyses — results, disaggregated by the declared factors.
  8. Ethical Considerations — risks and implications.
  9. Caveats and Recommendations — limitations and guidance.

Two things follow from the fact that this is a paper and not a standard. First, there is no conformance test: a document with two of the nine sections filled in is still, in ordinary usage, a model card. Second, no regulator anywhere mandates the format. It spread because model hosting platforms adopted it as a README convention, not because anyone required it.

Its sibling instrument, Datasheets for Datasets (Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III and Kate Crawford, arXiv:1803.09010, 23 March 2018), predates it by six months and applies the same logic one layer down, to the dataset: motivation, composition, collection process and recommended uses. Anyone who has written a data management plan will recognise the shape immediately.

The system card: what labs actually ship

A model card describes a model. A system card describes a deployment — the model plus the scaffolding it ships inside: safeguards, classifiers, refusal policies, system prompts, tooling, and the evaluation campaign run against the assembled thing rather than the raw weights.

This is the form frontier labs converged on. Anthropic’s transparency hub publishes a system card per model release, described as distilling “comprehensive technical assessments” covering capabilities and benchmarked performance, safety evaluation results across risk domains, training methodology, testing methods, deployment decisions, and alignment and safeguards assessments. OpenAI uses the same term for the same kind of document.

The distinction matters for anyone doing vendor diligence, because the two documents answer different questions. A model card tells you what the weights can do. A system card tells you what the product does once the vendor’s mitigations are in place — which is the only number that describes what your users will actually encounter. If you are assessing third-party AI vendor risk, a model card alone leaves the safeguards layer undescribed.

California SB 53 names both, and requires neither

California’s Transparency in Frontier Artificial Intelligence Act requires a frontier developer to publish a transparency report on or before the release of a new or substantially modified frontier model. The required contents are modest: the developer’s website, a mechanism for a natural person to contact the developer, the model’s release date, supported languages, output modalities, intended uses, and general restrictions or conditions on use. Large frontier developers add summaries of catastrophic-risk assessments conducted under their framework, the results, the extent of third-party evaluator involvement, and other steps taken to meet framework requirements for that model.

The statute then says, verbatim:

“A frontier developer that publishes the information described in paragraph (1) or (2) as part of a larger document, including a system card or model card, shall be deemed in compliance with the applicable paragraph.”

Read that carefully. SB 53 does not require a system card. It requires a list of facts, and says that if you happen to put them in a system card, that counts. The law is agnostic about the container. This is the closest any statute comes to endorsing either format, and it stops short of prescribing one. Our guide to California SB 53 covers the rest of the Act’s structure.

The EU Model Documentation Form: not a public document

The Model Documentation Form is the part of this picture most frequently misdescribed, usually as “the EU’s model card.” It is not public, and it is not a card.

It sits in the Transparency chapter of the General-Purpose AI Code of Practice, published by the European Commission on 10 July 2025. The Commission’s own description is blunt about what it is: the Transparency chapter “offers a user-friendly Model Documentation Form (DOCX) which allows providers to easily document the information necessary to comply with the AI Act obligation.” It is, literally, a Word template.

It exists to discharge two AI Act duties that are documentary rather than public. Article 53(1)(a) requires a provider to “draw up and keep up-to-date the technical documentation of the model, including its training and testing process and the results of its evaluation” — the detail of which is specified in Annex XI. Article 53(1)(b) requires a provider to “draw up, keep up-to-date and make available information and documentation to providers of AI systems who intend to integrate the general-purpose AI model into their AI systems” — specified in Annex XII.

The form covers licensing and distribution terms, model identification (name, version, release date, responsible entity), technical specifications including architecture, size and input/output formats, intended and approved use cases, integration dependencies across platforms and hardware, training methodology and design rationale, dataset details covering source, type, scope, volume, curation and bias mitigation, and energy and compute usage for both training and inference.

Its defining structural feature is the one that gets lost in summaries: the form tags each element by recipient. Some items are marked as relevant to the AI Office, some to national competent authorities, some to downstream providers who integrate the model. A single completed form therefore yields several different disclosures depending on who is asking. Documentation must be retained for each model version for ten years from the date that version was placed on the market, and provided on request — not proactively published.

Adherence to the Code of Practice is voluntary. Its function is evidentiary: providers can rely on it to demonstrate compliance with the underlying, binding Article 53 obligations. Signing is not the duty; Article 53 is. See the EU AI Act GPAI Code of Practice for who signed and what the other chapters cover.

The one that is genuinely mandatory: the public training-content summary

Article 53(1)(d) is the outlier. It requires every GPAI provider to “draw up and make publicly available a sufficiently detailed summary about the content used for training of the general-purpose AI model, according to a template provided by the AI Office.” The AI Office adopted that template on 24 July 2025, alongside an explanatory notice.

The template has three parts:

  • General information — identification of the provider, and of the model and its versions, including the base model where the release is a fine-tune of something else.
  • List of data sources — the main datasets used, and top domain names crawled.
  • Relevant data processing aspects — the legal aspects of the processing, sufficient for parties with a legitimate interest to exercise their rights under EU law, including copyright.

The obligation applied from 2 August 2025 for models placed on the market from that date. Models already on the market before then have until 2 August 2027 to publish a summary. Where a provider has made a genuine best effort but the information is unavailable or unreasonably burdensome to retrieve, the template requires the gap to be identified and explained rather than silently omitted — an explicit admission that incompleteness is expected.

Note the asymmetry this creates. The document that goes to regulators (the Model Documentation Form) is comprehensive and private. The document that goes to the public (the training-content summary) is narrow and mandatory. Neither is a model card. Neither has the nine sections.

The open-source carve-out splits them apart

Article 53(2) exempts models released under a free and open-source licence permitting access, use, modification and distribution, with publicly available parameters and architecture information — but not models designated as carrying systemic risk.

The exemption is partial, and the split runs exactly along the line described above. Article 53(1)(a) and 53(1)(b) fall away: an open-source provider owes no technical documentation to the AI Office and none to downstream providers. Article 53(1)(c), the copyright policy, and Article 53(1)(d), the public training-content summary, both survive. An open-weight release therefore escapes the Model Documentation Form entirely while still owing the public summary. See what Article 53’s open-source exemption actually carves out, and GPAI systemic risk for the designation that removes the carve-out altogether.

Where this reaches research administration

This is not only a frontier-lab problem. Three parts of it land on university research offices.

Releasing a model can make an institution a provider. A GPAI model developed solely for scientific research and development does not make its developer a provider under the AI Act. But that exemption attaches to the purpose, not to the institution: a model a research group releases beyond that narrow research context does not automatically inherit it. Nonprofit or academic status is not itself a defence. A lab that fine-tunes a base model and pushes the weights to a public hub has made a release decision with potential Article 53 consequences, and in most universities nobody in sponsored programs or research computing is in that approval path.

Fine-tuning has a quantified threshold. A downstream modifier becomes a provider in its own right when the compute used for the modification exceeds one third of the original model’s training compute. Where the base model’s compute is unknown, the Commission’s guidance offers fallbacks: one third of 1023 FLOP for a general-purpose model, one third of 1025 FLOP for a systemic-risk model. Above the line, the obligations attach, but only to the modifications made. Ordinary academic fine-tuning sits far below these numbers — which is the useful finding, and worth recording once rather than re-litigating per project.

Documentation practice already exists here. Datasheets for datasets, data management plans, provenance records and IRB documentation are the same instinct applied to research data. An institution asked to produce a training-content summary for a model trained on its own corpora will find the raw material in files sponsored programs and the data office already hold. The gap is usually organisational, not evidentiary.

Where NIKOLAI fits

NIKOLAI is CASRAI’s own independent frontier-AI-safety dictionary. It is unendorsed: no lab, evaluator or regulator has adopted it, and its crosswalk rows are shadow mappings — CASRAI’s reading of published documents, not something the named organisations have agreed to.

The gap NIKOLAI’s track N8, Transparency and review, points at here is specific. All four documents above describe what is present. None has a field for what was taken out. A system card that omits a capability evaluation, and a Model Documentation Form that withholds an item as a trade secret, look identical to a complete one from the outside.

NIKOLAI’s proposed redaction element is a structured record of a removal from a published artefact, capturing where the removal occurred, how much was removed, why (from a controlled redaction-reason vocabulary), who decided, and — the load-bearing part — a disclosure that the removal happened at all. That last field is what turns an absence into a fact a reader can act on. Our guide to what gets redacted from AI safety reports works through the permitted and prohibited reasons in more depth.

Frequently asked questions

Is a model card legally required anywhere?

No. No jurisdiction mandates the model card format. California SB 53 mentions model cards and system cards only as acceptable containers for a list of facts it separately requires; the EU AI Act requires documentation and a public training-content summary, neither of which follows the nine-section model card structure.

What is the difference between a model card and a system card?

A model card documents the model: architecture, training data, intended use, disaggregated performance. A system card documents the deployed system: the model plus safeguards, classifiers, policies and tooling, evaluated as assembled. Frontier labs publish system cards because the safeguards layer materially changes what the product does.

Is the EU Model Documentation Form published?

No. It is completed and retained by the provider, then supplied on request to the AI Office, national competent authorities or downstream providers, with items tagged by recipient. Retention is ten years per model version from the date it was placed on the market. The document the public sees is the separate Article 53(1)(d) training-content summary.

Does the open-source exemption remove the documentation duty?

Partly. Article 53(2) removes 53(1)(a) technical documentation and 53(1)(b) downstream documentation for qualifying free and open-source releases. The 53(1)(c) copyright policy and the 53(1)(d) public training-content summary both remain. Models designated as carrying systemic risk get no exemption at all.

Does a university that releases a fine-tuned model take on AI Act obligations?

Possibly. The scientific research and development exemption covers models developed solely for that purpose; a release beyond that context does not automatically inherit it. A modifier becomes a provider when modification compute exceeds one third of the base model’s training compute, with fallbacks of one third of 1023 FLOP generally and one third of 1025 FLOP for systemic-risk models. Typical academic fine-tuning falls well below both.

When does the training-content summary have to be published?

From 2 August 2025 for models placed on the EU market on or after that date. Models already on the market before then have until 2 August 2027.

Sources

Related reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about Model Cards, System Cards, and the EU Model Documentation Form

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.