Skip to main content
v2026.11,858 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

BIG-bench

The Beyond the Imitation Game benchmark, a community-contributed collection of more than 200 tasks designed to probe capabilities of large language models that may be missed by narrower benchmarks.

ByCASRAI Editorial Board
· Last updated 5 Sept 2026
Share this

Ask CASRAI · free to try

Ask about BIG-bench

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    A model technical report including BIG-bench Hard average accuracy across 23 tasks.

  • Is an instance

    A new benchmark paper using BIG-bench as a baseline distribution of LLM capability.

Counter-examples

Looks similar, but isn't

  • Not an instance

    MMLU (a different benchmark).

  • Not an instance

    A single dataset like SQuAD.

Editorial commentary

BIG-bench (the Beyond the Imitation Game benchmark) is a community-contributed collection of more than 200 tasks, published in 2022 (Srivastava et al., arXiv:2206.04615), designed to probe language-model capabilities — reasoning, world knowledge, social bias, and more — that narrower, single-metric benchmarks tend to miss. Tasks were submitted and peer-reviewed by contributors across many institutions rather than authored by one lab.

Currency note: this benchmark is now largely historical

BIG-bench saturated as models improved: state-of-the-art systems now score near-ceiling on most of its original tasks, which is why a harder subset, BIG-Bench Hard (BBH, Suzgun et al., 2022), was carved out specifically from the tasks where models still underperformed humans. As of 2026, BBH itself has substantially saturated too — current frontier models score well above 0.9 on the public BBH leaderboard — prompting a further successor, BIG-Bench Extra Hard (BBEH, 2025), built by replacing each BBH task with a harder variant probing the same underlying reasoning capability; best-performing models on BBEH score roughly in the 10-45% range as of its introduction, restoring the discriminative power the original benchmark has lost. A page, procurement document, or grant proposal citing a raw BIG-bench (not BBH or BBEH) score as evidence of current model capability should be read with this saturation in mind.

How this differs from its siblings

See MLCommons benchmark for the fixed-workload, hardware-comparison alternative, and synthetic benchmark for benchmarks whose items are model-generated rather than human-authored (BIG-bench’s tasks were human-authored and peer-reviewed).

References

  • Srivastava, A. et al. (2022). “Beyond the Imitation Game.” arXiv:2206.04615.
  • Suzgun, M. et al. (2022). “Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.”
  • “BIG-Bench Extra Hard” (2025), arXiv:2502.19187.
  • See also: MMLU benchmark, HELM benchmark.

Also known as

Beyond the Imitation Game Benchmark · BBH (subset)

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="BIG-bench"
      vocab-term-identifier="https://casrai.org/dictionary/term/big-bench" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/big-bench",
  "name": "BIG-bench",
  "identifier": "https://casrai.org/dictionary/term/big-bench",
  "description": "The Beyond the Imitation Game benchmark, a community-contributed collection of more than 200 tasks designed to probe capabilities of large language models that may be missed by narrower benchmarks.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/big-bench",
  "alternateName": [
    "Beyond the Imitation Game Benchmark",
    "BBH (subset)"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:51",
  "dateModified": "2026-09-05T14:24:39",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →