Skip to main content
v2026.11,610 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

Red-teaming

The practice of deliberately adversarial testing of an AI system by skilled testers attempting to elicit failures, unsafe outputs, or policy violations, in order to discover weaknesses before deployment.

ByCASRAI Editorial Board
· Last updated 21 May 2026

Ask about Red-teaming

Answers are drawn from this dictionary entry and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Examples

Worked examples

  • Is an instance

    A frontier-lab internal red team running 6 weeks of structured probes pre-deployment and recording all elicited failures.

  • Is an instance

    A public red-team event at DEF CON 31 inviting 2,200 participants to probe four frontier models.

Counter-examples

Looks similar, but isn't

  • Not an instance

    Standard benchmark evaluation.

  • Not an instance

    User feedback gathered post-deployment.

Editorial commentary

AI red-teaming draws on traditions from information-security and military adversarial testing. It complements automated evaluation by uncovering issues that benchmarks miss. Red-team reports are increasingly disclosed alongside frontier models (OpenAI GPT-4, Anthropic Claude, Meta Llama). DEF CON 31's Generative Red Team event normalised public red-team exercises.

References

  • Ganguli et al., 'Red Teaming Language Models' (arXiv 2022); OpenAI 'GPT-4 System Card' (2023).

Also known as

AI red team · adversarial AI testing

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Red-teaming"
      vocab-term-identifier="https://casrai.org/dictionary/term/red-teaming" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/red-teaming",
  "name": "Red-teaming",
  "identifier": "https://casrai.org/dictionary/term/red-teaming",
  "description": "The practice of deliberately adversarial testing of an AI system by skilled testers attempting to elicit failures, unsafe outputs, or policy violations, in order to discover weaknesses before deployment.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/red-teaming",
  "sameAs": [
    "AI red team",
    "adversarial AI testing"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "dateModified": "2026-05-21T02:22:51",
  "inLanguage": "en"
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →