Skip to main content
v2026.11,858 entries · CC-BY 4.0

Third-Party AI Evaluator Standards: Independence, Access, and Methodology

A cross-cutting guide to the standards questions that apply to every third-party AI evaluator: embedded vs. arms-length access, independence and conflict-of-interest, evaluation validity, red-team access agreements, publication rights, and retaliation protection — mapped to NIKOLAI’s N8 Transparency and Review track.

Written and maintained by CASRAI Editorial Board

Last updated

Why methodology, not just access, is the open question

Frontier AI labs now grant a growing list of outside groups some form of access to their models before or after release: government testing institutes such as the US CAISI and UK AISI, independent nonprofits such as METR, and compliance auditors working against frameworks such as ISO/IEC 42001 or the EU AI Act. Each of those groups is a different kind of evaluator, and CASRAI’s guides to CAISI and UK AISI’s pre-deployment testing, METR, and third-party AI auditing cover what each one is and does.

This guide sits a layer above those institution-specific pages. It is about the standards and methodology questions that apply across the whole third-party evaluator ecosystem, whichever organisation is doing the evaluating: what “independent” is supposed to mean, what access an evaluator actually needs, what can go wrong with the evaluation itself, and what happens after the evaluator produces a finding the developer would rather not see published.

Embedded vs. arms-length evaluator models

Third-party evaluation happens under two broadly different working arrangements, and the difference shapes almost everything else in this guide.

  • Arms-length evaluation. The evaluator receives a model (often via API, sometimes with elevated access) for a defined testing window, runs its evaluation suite, and reports findings back. METR’s published evaluations and most compliance audits against ISO/IEC 42001 or the EU AI Act’s Article 43 conformity-assessment route follow this pattern: a bounded engagement with a deliverable at the end.
  • Embedded evaluation. The evaluator has sustained access during model development itself — sitting alongside the developer’s own safety testing rather than receiving a finished model afterward. CASRAI’s guide to CAISI and UK AISI’s pre-deployment testing describes this as an emerging, still-informal term for the kind of ongoing access some government evaluators have negotiated with frontier labs.

Embedded access can surface problems earlier, before a model ships, but it also raises the independence questions in the next section more sharply: an evaluator working inside the development process for months is closer to the people and incentives it is supposed to be assessing than one that receives a model for a two-week test window.

Evaluator independence and conflict-of-interest standards

“Independent” is doing a lot of work in the phrase “independent third-party evaluator,” and it is not self-enforcing. The recurring independence questions across evaluator types are:

  • Funding source. Does the evaluator take money from the company whose systems it evaluates? METR states it has not accepted funding from AI companies and treats this as a condition of being able to evaluate their models independently. Government evaluators such as CAISI and UK AISI are publicly funded rather than lab-funded, which removes this specific conflict but introduces a different one — political and diplomatic pressure on a state body.
  • Personnel and prior relationships. Has the evaluator, or the specific staff assigned to an engagement, previously worked for, consulted to, or held equity in the company being evaluated?
  • Editorial control over findings. Can the evaluated company edit, delay, or veto the evaluator’s report before publication? This overlaps with publication rights, covered below, but the independence question is narrower: it asks whether the evaluator’s conclusions are its own, not whether they get published at all.
  • Scope-setting. Who decides what gets tested? An evaluator that can only test what the developer proposes has a narrower kind of independence than one that sets its own scope.

CASRAI’s NIKOLAI vocabulary formalises this as a specific, named element: Evaluator Independence and Conflict of Interest (N8.1), part of NIKOLAI’s Transparency and Review track. Rather than leaving “independent” as an adjective in a press release, N8.1 gives organisations a defined field for declaring the ties between an evaluator and the developer it is assessing, so that independence claims can be checked rather than taken on faith.

Evaluation-validity threats and benchmark saturation

A conflict-free evaluator can still produce a misleading evaluation if the evaluation itself is a poor measure of the thing it claims to measure. Two validity threats come up repeatedly in frontier AI evaluation:

  • Benchmark saturation. When a large share of models score near the maximum on a given benchmark, that benchmark stops usefully distinguishing between them — and stops providing much signal about further capability gains. This is a known limitation of static, fixed benchmarks generally, not specific to any one test, and it is a reason evaluators such as METR report continuous metrics (like the task-completion time horizon described in CASRAI’s METR guide) rather than relying solely on percentage-correct scores against a fixed question set.
  • Elicitation quality. A model’s measured capability depends heavily on how hard the evaluator tried to elicit it — prompting strategy, scaffolding, fine-tuning, and compute budget all affect the result. An evaluation that under-elicits a capability can understate real-world risk just as badly as a saturated benchmark can overstate discriminating power. CASRAI’s METR guide notes that METR documents its elicitation methodology explicitly for this reason.

Neither problem has a single agreed fix. The practical implication for anyone reading an evaluation report is to check what the evaluator says about how it elicited the capability being measured and whether the benchmark it used still separates strong models from weaker ones, rather than treating a headline score as self-explanatory.

Red-team access agreements for external evaluators

An evaluator can only find what it is allowed to look for. The access an external evaluator is granted — API-only access, fine-tuning access, access to model internals or training data, or red-team access to probe for adversarial failure modes — determines the ceiling on what its evaluation can credibly claim. A report based on default, safety-filtered API access cannot support the same conclusions as one based on an agreement granting unrestricted red-team access with safety mitigations removed.

Because the access an evaluator was actually given is rarely visible from the published report alone, NIKOLAI defines a dedicated element for it: Evaluator Access Attestation (N8.4), which documents what access was granted or withheld, and any confidentiality terms attached to it, as a discrete, checkable record rather than leaving it implicit in the evaluator’s prose.

Publication rights: can an evaluator publish an inconvenient finding?

The independence and access questions above matter less if the evaluator cannot publish what it finds. A publication rights clause in an evaluator–developer agreement typically covers who decides whether a finding is published at all, whether the developer gets pre-publication review or veto rights, what the developer can require be redacted (commonly citing safety or security grounds), and what happens if the evaluator and developer disagree about a redaction request.

NIKOLAI’s Transparency and Review track treats this as its own element, Publication Rights Clause (N8.3), and pairs it with two further elements for the case where content is withheld: Redaction (N8.7) records that something was removed, and Redaction reason (N8.6) records why. The track’s own stated rationale is direct about what is at stake: external review “only means something if the evaluator is independent and a developer can’t quietly redact an inconvenient finding” — which is why redaction is modelled as a visible, reasoned event rather than a silent omission.

Protecting evaluators from retaliation and funding pressure

Publication rights address what happens to a specific finding. A related but separate question is what happens to the evaluator itself after it publishes something the developer did not want published: whether continued access to future models is withdrawn, whether funding or contracts are reduced, or whether the evaluator faces less formal reputational pressure. An evaluator that reasonably expects to lose access or funding after an unfavourable finding faces a structural incentive to soften its conclusions, regardless of how the independence and publication-rights terms are written on paper.

This is the least standardised part of third-party evaluation practice. Unlike independence, access, and publication rights, it does not yet have a widely adopted named clause or NIKOLAI element of its own; where it is addressed at all today, it tends to be handled informally, through an evaluator’s funding diversification (as with METR’s donor base) or through a government evaluator’s statutory footing rather than through a specific contractual anti-retaliation provision. It is worth asking about explicitly in any evaluator–developer agreement precisely because it is not yet standard practice to include it.

Where NIKOLAI formalises this: the Transparency and Review track

The elements referenced throughout this guide — evaluator independence, access attestation, publication rights, and redaction — are not CASRAI’s informal commentary on the field. They are drawn directly from NIKOLAI’s N8 (Transparency and Review) track, the part of the NIKOLAI dictionary that defines the vocabulary organisations use to document how an AI system was reviewed and by whom. The full N8 track has seven elements: Evaluator Independence and Conflict of Interest (N8.1), AI-Model Review (N8.2), Publication Rights Clause (N8.3), Evaluator Access Attestation (N8.4), External Review (N8.5), Redaction reason (N8.6), and Redaction (N8.7).

Read together, N8.2 (AI-Model Review) and N8.5 (External Review) capture that a review happened and what it covered, while N8.1, N8.3, N8.4, N8.6, and N8.7 capture the conditions that determine whether that review’s findings can be trusted: who ran it, what they were allowed to see, and whether anything was left out and why. This guide’s structure follows that same division deliberately — it is, in effect, the explanatory companion to what N8 already defines formally.

Frequently asked questions

Is an “embedded” evaluator less independent than an arms-length one?

Not automatically, but the risk is structurally higher. Sustained, in-development access can produce earlier and more thorough findings, but it also means the evaluator spends more time inside the developer’s own processes and relationships, which is exactly what independence standards such as NIKOLAI’s N8.1 are designed to make visible and checkable rather than assumed.

Does a conflict-of-interest disclosure make an evaluation independent?

Disclosure and independence are related but not the same thing. A disclosed conflict is still a conflict; NIKOLAI’s N8.1 element is a record of declared ties, not a certification that none exist. Readers of an evaluation report still need to weigh what a disclosed relationship means for the specific findings involved.

Who decides what an evaluator is allowed to test?

It varies by agreement. In some arrangements the developer sets the scope and grants access accordingly; in others, particularly with government evaluators operating under a testing agreement, the evaluator has more say in what it tests for. NIKOLAI’s Evaluator Access Attestation (N8.4) is designed to make the actual access granted — whoever decided it — a matter of record rather than something inferred from the published report.

Can a developer legally block an evaluator from publishing a finding?

This depends entirely on the underlying agreement; there is no general legal right to publish that overrides a contract’s terms. That is precisely why publication rights are treated as a specific, negotiated clause — formalised in NIKOLAI as the Publication Rights Clause (N8.3) — rather than assumed to follow automatically from an evaluator’s independence.

How is this guide different from CASRAI’s guides on METR, CAISI/UK AISI, or third-party AI auditing?

Those guides are about specific organisations and what each one does. This guide is about the standards and methodology questions — independence, access, validity, publication rights — that apply across all of them, whichever evaluator is involved. Read the institution guides for what a given evaluator does; read this one for how to judge whether any evaluator’s findings can be trusted.

This guide is part of CASRAI’s Third-Party Evaluation & Assurance subcluster within the Frontier AI Safety & Governance content cluster. For the full vocabulary this guide draws on, see the NIKOLAI dictionary.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about Third-Party AI Evaluator Standards: Independence, Access, and Methodology

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →