Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §1.3.5
“Version 3.2 of our RSP authorizes Anthropic's Long-Term Benefit Trust (LTBT) to request external review of risk reports, gives the LTBT the power to approve our selection of external reviewers ... every part of any unredacted risk report is reviewed by at least one such reviewer.”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfEU GPAI Code of Practice, Safety and Security Chapter, Appendix 3.5; Measures 7.2, 7.4
“Signatories will ensure that adequately qualified independent external evaluators conduct model evaluations pursuant to this Appendix 3, with regards to the systemic risk, unless the model is exempt as 'similarly safe or safer' or Signatories cannot find qualified evaluators 'despite using early search efforts (such as through a public call open for 20 business days).'”
https://ec.europa.eu/newsroom/dae/redirection/document/118119METR, accountability element
“External scrutiny ensures that a company's safety claims can be independently validated by qualified experts.”
https://metr.org/
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report, August 2026 | “Version 3.2 of our RSP authorizes Anthropic's Long-Term Benefit Trust (LTBT) to request external review of risk reports, gives the LTBT the power to approve our selection of external reviewers ... every part of any unredacted risk report is reviewed by at least one such reviewer. The LTBT has not requested an external review (nor has the RSP required it), though pilot external reviews continue.” Also draws on the Advanced AI Framework's evaluator-review contents (AAF p.8) {AAF}. | exact confidence: high | Anthropic Risk Report, August 2026 |
| OpenAI OpenAI Preparedness Framework v2 / Frontier Governance Framework | “"then when available and feasible, OpenAI will work with third-parties to independently evaluate models" (PF §5.2); "We may solicit and obtain input from external experts" (FGF §5).” Also cites Frontier Governance Framework §5 {FGF}. | close confidence: medium | OpenAI Preparedness Framework v2 |
| Google DeepMind Gemini 3.7 Flash FSF report | “"External safety testing" by "specialist independent groups"; "The independent external evaluators explore frontier safety risk domains, however they do not directly comment on risk thresholds as contemplated under our FSF." (p.6)” | close confidence: medium | Gemini 3.7 Flash FSF report |
| EU EU GPAI Code of Practice, Safety and Security Chapter | “"Signatories will ensure that adequately qualified independent external evaluators conduct model evaluations pursuant to this Appendix 3, with regards to the systemic risk", unless exempt as 'similarly safe or safer' or unable to find qualified evaluators 'despite using early search efforts.' Model Report must explain 'the choice of evaluator based on the qualification criteria' and whether independent evaluator input 'informed' the decision to proceed (Measures 7.2, 7.4).” | exact confidence: high | EU GPAI Code of Practice, Safety and Security Chapter |
| California SB 53 California SB 53 | “Framework topic "(5) Using third parties to assess the potential for catastrophic risks and the effectiveness of mitigations of catastrophic risks"; transparency report summary of "(C) The extent to which third-party evaluators were involved."” | close confidence: medium | California SB 53 |
| US Government (NIST CAISI) / UK AI Safety Institute (AISI) partnership NIST CAISI bulletin; DeepMind UK AISI partnership post | “CAISI: "CAISI will conduct pre-deployment evaluations and targeted research to better assess frontier AI capabilities." UK AISI with DeepMind: "partnered with the UK AISI since its inception in November 2023 to test our most capable models."” Also cites {AISI}. | close confidence: medium | NIST CAISI bulletin |
| METR METR (site) | “"External scrutiny ensures that a company's safety claims can be independently validated by qualified experts."” | exact confidence: high | METR |
| Frontier Model Forum FMF Third-Party Assessments technical report | “"Third-party assessment: 'Third-party assessments can be conducted on frontier models to confirm evaluations or claims on critical safety capabilities and mitigations.'" Functions: "Confirmation", "Robustness", "Supplementation".” | exact confidence: high | FMF Third-Party Assessments |
| G42 G42 Frontier Safety Framework | “"G42 will engage in annual external audits to verify compliance with the Framework" (s.5).” | close confidence: medium | G42 Frontier Safety Framework |
| Anthropic Policy on the AI Exponential (Amodei essay) | “"Models above a threshold of compute should undergo mandatory testing by a qualified third party for their level of risk in four specific areas: cybersecurity, biological weapons, loss of control of AI systems, and automated R&D." Delivery left as a choice between "a government agency (similar to the FAA)" or "a set of private organizations that are authorized and inspected by the government ... (a 'regulatory markets' approach)."” Proposes a distinct, two-track mandatory external-review architecture separate from the voluntary/embedded-evaluator model in {PACE}; neither track specifies evaluator access level, publication rights or independence criteria -- a gap the essay itself does not fill. | close confidence: medium | Policy on the AI Exponential |
| US Congress (H.R. 9925, FRONTIER Act, not enacted) FRONTIER Act, 119th Congress | “"A very large frontier developer shall grant an IVO timely access upon request to unredacted materials, records, personnel, systems, and all other information reasonably necessary for conducting the assessments," with assessments "not less frequently than once every 6 months" (Sec. 5). Large frontier developers separately "commission an annual independent compliance audit" that a named "lead auditor" signs to certify (Sec. 4(c)).” Richest independent-verification-organization licensing design in the corpus, but ties to a mandatory-audit model rather than an embedded-evaluator or publication-rights framing: IVO access is 'upon request,' not standing employee-like access, and the bill 'does not settle publication rights for assessors beyond the redacted public copy.' Not enacted. | close confidence: medium | FRONTIER Act, 119th Congress (H.R. 9925, introduced, not enacted) |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- xAI
"There is no commitment to external or third-party evaluation" (brief finding, June 2026). Cards separately state "third-party testing shows that Grok 4's end-to-end offensive cyber capabilities remain below the level of a human professional" and that third-party evaluators "corroborated" Grok 4.6 cyber results -- RL, a pointer not a mapping.
xAI Frontier AI Framework, 30 Jun 2026 - Meta
External experts used in threat-modeling workshops, risk assessments "where appropriate" and red teaming "when appropriate" -- RL, a pointer not a mapping.
Meta Advanced AI Scaling Framework v2







