Skip to main content
v2026.11,858 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

AI evaluation card

A structured documentation artefact specifically describing an evaluation of an AI system, separate from the model card, including the evaluation methodology, datasets, metrics, results, and known limitations of the evaluation itself.

ByCASRAI Editorial Board
· Last updated 5 Sept 2026
Share this

Ask CASRAI · free to try

Ask about AI evaluation card

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    An evaluation card for a code-completion benchmark documenting the held-out test set, prompting template, and decoding configuration.

  • Is an instance

    An evaluation card for a clinical-reasoning probe describing rater calibration.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A model card section labelled 'Evaluation' but not separately documented.

  • Not an instance

    A leaderboard table without supporting methodology.

Editorial commentary

An AI evaluation card is a structured documentation artefact that describes an evaluation of an AI system — the evaluation methodology, the datasets used, the metrics reported, the results, and the known limitations of the evaluation itself — kept separate from the model card that documents the system being evaluated.

How this differs from the other AI-governance documentation and process terms

An evaluation card documents a specific act of testing: what was measured, how, against what benchmark, with what caveats. It is not the same artefact as an AI conformance assessment, which is a formal, often regulatory, determination that a system meets applicable legal or standards-based requirements before market placement — a conformance assessment may draw on evaluation-card-style evidence, but it is a compliance judgment, not a documentation format. It is also distinct from a model card (which documents the system’s design, training data, and intended use) and from a ISO/IEC 42001 AI management system (which governs an organisation’s overall AI processes, not a single evaluation).

Why evaluation itself needed its own documentation format

Evaluation cards are a relatively recent (2023 onward) addition to the documentation-artefact family, reflecting the recognition that an evaluation — a dataset, a protocol, an analysis approach — is itself a reusable artefact deserving its own documentation, separate from documenting the model being evaluated. NIST’s GenAI evaluation profile work and Stanford’s Center for Research on Foundation Models (CRFM) evaluation reports are commonly cited examples of the genre, each making explicit what a given benchmark result does and does not demonstrate.

Sources

NIST AI RMF Generative AI Profile; Stanford CRFM evaluation transparency reports.

Also known as

eval card

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="AI evaluation card"
      vocab-term-identifier="https://casrai.org/dictionary/term/ai-evaluation-card" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/ai-evaluation-card",
  "name": "AI evaluation card",
  "identifier": "https://casrai.org/dictionary/term/ai-evaluation-card",
  "description": "A structured documentation artefact specifically describing an evaluation of an AI system, separate from the model card, including the evaluation methodology, datasets, metrics, results, and known limitations of the evaluation itself.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/ai-evaluation-card",
  "alternateName": [
    "eval card"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:51",
  "dateModified": "2026-09-05T14:24:50",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.