Skip to main content
v2026.11,858 entries · CC-BY 4.0

Direct comparison

CAISI vs NTIA: Open-Weight AI Playbooks

CAISI tests released open-weight models for cyber risk. NTIA's 2024 report sets the policy logic for restricting release at all.

Written and maintained by CASRAI Editorial Board

Last updated

Ask CASRAI · free to try

Ask about CAISI vs NTIA: Open-Weight AI Playbooks

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

How do CAISI (NIST), NTIA compare side by side?

The table below compares CAISI (NIST), NTIA across 7 procurement-relevant dimensions, from what kind of document through binding force.

Side-by-side comparison

DimensionCAISI (NIST)NTIA
What kind of documentA technical cyber-capability assessment of one named model at a time — GLM-5.2 (published July 17, 2026), Kimi K3 (preliminary joint assessment with UK AISI, July 23, 2026, updated August 28, 2026), and GLM-5.3 (September 17, 2026).A single policy report, “Dual-Use Foundation Models with Widely Available Model Weights”, published July 30, 2024 in response to Section 4.6 of Executive Order 14110. It evaluates no specific model.
When it acts, relative to releaseAfter the fact. Each assessment begins only once a model's weights are already public — CAISI's own cadence has run four to five weeks from a model's release date to its published assessment in every case so far.Before the fact, in principle: it is a standing framework meant to inform decisions about future models and future releases, not a review triggered by any one model's launch.
The question it is answering“How capable is this specific, already-released open-weight model at cyber offense, measured against a repeatable benchmark suite?”“Should the US government restrict who may release an open-weight model's weights at all, now or later — and if so, on what evidence?”
ScopeNarrow and technical: cyber-capability only, via benchmark families such as ExploitBench, SEC-Bench Pro, CAISI OSS-Fuzz, ExploitGym Userspace, and the “The Last Ones” cyber range. It does not rule on release policy.Broad and structural: covers marginal risk, benefits (research access, competition, transparency), and the mechanics of restriction itself — export controls, licensing, liability — not any single capability domain.
What it actually concludesA capability estimate for one model at one point in time — e.g. GLM-5.2 assessed as roughly comparable in cyber capability to Claude Opus 4.6, with safeguards that CAISI judged did not reliably block agentic exploit assistance once self-hosted.A deliberate non-conclusion: “current evidence is not sufficient to definitively determine either that restrictions on such open weight models are warranted, or that restrictions will never be appropriate in the future.” It declines to restrict release today.
What it proposes going forwardA repeatable evaluation cadence, not a policy recommendation — CAISI's role stops at publishing capability findings; it does not itself decide whether a model's release should have been restricted.A four-part adaptive framework: (1) collect evidence through audits, disclosures, and risk research; (2) evaluate that evidence, including the capability gap between open and closed models; (3) act — potential access restrictions or other mitigations — if warranted; (4) maintain flexibility for further government action as circumstances change.
Binding forceNone. CAISI's assessments are findings, not enforceable determinations — no US law currently restricts publishing an open-weight model's cyber-capability level or requires pre-release review of it.None. The NTIA report is Commerce Department guidance to inform future policymaking, not a rule, statute, or licensing requirement. It creates no new restriction and imposes no new duty on anyone.

Common questions

Common questions about CAISI (NIST) vs NTIA

Does NTIA's report cover the same models CAISI assessed?

+

No — they can't, on the dates involved. NTIA's report was published July 30, 2024, more than a year and a half before GLM-5.2 (June 2026), Kimi K3 (July 2026), or GLM-5.3 (August 2026) existed. NTIA's report is a standing policy framework meant to apply to open-weight models generally, including ones released long after it, not a review of any named model.

Did NTIA recommend restricting open-weight AI models?

+

No. The report's own language is explicit that it does neither: “current evidence is not sufficient to definitively determine either that restrictions on such open weight models are warranted, or that restrictions will never be appropriate in the future.” It sets up a process for revisiting that question — collect evidence, evaluate it, act if warranted, stay flexible — rather than answering it either way in 2024.

Does CAISI's work feed into NTIA's framework, or the other way around?

+

Neither directly, on the public record. CAISI (created within NIST in 2025) and NTIA are both Commerce Department components, and CAISI's per-model cyber assessments are exactly the kind of technical “evidence collection” NTIA's 2024 framework calls for in its first phase — but CAISI's assessments themselves don't cite NTIA's report, and NTIA's report predates CAISI's 2026 assessment cadence by nearly two years. Treat them as two separate outputs of the same department's broader open-weight-model posture, not a documented pipeline from one to the other.

Does NIKOLAI have a crosswalk row for NTIA's report?

+

No, and this page doesn't claim one exists. NIKOLAI's N1 track defines a Coverage scope threshold element — the if-then test deciding whether a framework applies to a given model — and its live crosswalk maps eleven labs, statutes, and proposals (Anthropic, OpenAI, the EU AI Act, SB 53, the FRONTIER Act, and others) against that element. None of those rows is NTIA's report, and none currently exists for it: NTIA's framework sets a process for deciding whether to regulate, rather than a scope test with a specific compute or revenue threshold that a NIKOLAI crosswalk row is built to record.

Is there a genuine NIKOLAI tie-in for CAISI's side of this comparison?

+

Yes — a real, live one. NIKOLAI's N1 track includes a Risk domain element (the top-level category of catastrophic harm a framework's evaluations are organized around), and its crosswalk already carries a shadow-mapped row for the US Government/CAISI reading it as narrow-scope: “advanced cyber capabilities” only — unlike Anthropic, OpenAI, or the EU AI Act, which each name three to four domains (CBRN, cyber, loss of control, harmful manipulation). That's CASRAI's own independent, unendorsed reading of CAISI's public assessments, not a mapping CAISI has filed or endorsed itself — but it is a specific, checkable claim, not an invented one.

Why does this matter for research administration?

+

University labs that train and openly release model weights already route that release decision through their institution's export-control/research-security office before publication — not through NIST or NTIA. That review turns on the same public-availability logic NTIA's 2024 report examines: technology and software that qualify as "published" or "publicly available" generally sit outside export-licensing requirements, and research-security offices apply the closely related Fundamental Research Exclusion when clearing a lab's intent to publish code, results, or model weights openly. A research-security officer deciding whether to greenlight release of a university-trained open-weight model is, at smaller scale, making the same call NTIA's framework poses.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →