Written and maintained by CASRAI Editorial Board
Last updated
On September 17, 2026, NIST’s Center for AI Standards and Innovation (CAISI) published its assessment of Z.ai’s GLM-5.3 — the model Z.ai had released just over a month earlier, on August 14, 2026. It was not CAISI’s first look at a Chinese-lab open-weight model, and it will not be its last: over the four months before it, CAISI (sometimes jointly with the UK AI Security Institute, UK AISI) had already published a cyber-capability assessment of Z.ai’s GLM-5.2 (July 17, 2026) and a preliminary joint assessment of Moonshot AI’s Kimi K3 (July 23, 2026). That is a real, dated, primary-sourced cadence — not a single reactive report — and it answers a governance question that is genuinely distinct from the closed-frontier-lab safety-framework content that dominates most AI-governance coverage: what happens when a highly capable model’s weights are simply published, so there is no API to gate and no deployment decision to review?
This guide covers what CAISI has actually assessed so far, how its evaluation methodology works across those assessments, and why open-weight, PRC-origin model risk is its own governance category. For CAISI as an institution — its mandate, its 2025 rename, and its relationship to the UK AI Security Institute — see our companion guide. For the mechanics of CAISI’s pre-deployment testing agreements with individual labs, see our guide to those agreements.
The cadence so far
Three assessments, each targeting a newly released open-weight model from a PRC-based lab, published in close succession:
| Published | Model | Developer | Model released | Collaboration |
|---|---|---|---|---|
| July 17, 2026 | GLM-5.2 | Z.ai (formerly Zhipu AI) | June 16, 2026 | CAISI |
| July 23, 2026 | Kimi K3 | Moonshot AI | July 16, 2026 | UK AISI & CAISI (preliminary joint assessment, updated August 28, 2026) |
| September 17, 2026 | GLM-5.3 | Z.ai | August 14, 2026 | CAISI |
In each case, CAISI (and UK AISI, for the joint Kimi K3 assessment) moved from a model’s public release to a published cyber-capability assessment in roughly four to five weeks. That turnaround, repeated three times against three different labs’ releases inside one quarter, is the actual evidence of an ongoing program rather than a one-off response to a single model.
How CAISI evaluates open-weight cyber capability
Across these assessments, CAISI’s approach centers on aggregating performance across many individual tasks into a single capability estimate, rather than reporting raw pass rates in isolation. The GLM-5.3 and Kimi K3 assessments both describe an approach built on (or “inspired by”) Item Response Theory (IRT), producing a cyber capability index in which a fixed-size increase in score corresponds to a fixed multiplicative increase in the odds of solving a task — a scale designed to let capability be compared across models and over time, not just within one report.
The specific tasks vary by assessment but cluster around a few benchmark families:
- ExploitBench — a public benchmark (Carnegie Mellon University) measuring how far a model can progress along the software-exploitation ladder against a fixed set of post-2023 V8 JavaScript-engine vulnerabilities.
- SEC-Bench Pro and CAISI OSS-Fuzz — locating and exploiting vulnerabilities in browser engines (V8, SpiderMonkey) and in widely used open-source projects, including a large set of non-public tasks CAISI maintains specifically so models cannot have trained on the answers.
- ExploitGym Userspace — developing exploits from bugs in open-source projects at larger scale.
- “The Last Ones” (TLO) cyber range — a simulated 32-step corporate-network attack path across roughly 20 hosts and 4 subnets, used in the Kimi K3 assessment to measure sustained, multi-stage offensive capability rather than single-exploit development.
Models are run inside an agent harness (a ReAct-style loop with bash and Python tools, turn limits in the low hundreds, maximum reasoning settings enabled) rather than queried for one-shot answers — the assessments are explicitly measuring what an agentic system built on the model can accomplish, not what the base model can describe. For the closed U.S. models used as comparators, CAISI evaluates with system-level safeguards disabled, on the stated basis that this is necessary to measure maximal capability rather than a lab’s current refusal behavior; the open-weight models under assessment are, by construction, already runnable without a vendor’s safeguards once self-hosted.
What the three assessments found
GLM-5.2 (July 2026): CAISI’s assessment describes GLM-5.2 as “probably the most capable open-weight AI model when it was released,” with overall capability comparable to GPT-5.2 (released December 2025) and cyber capability comparable to Claude Opus 4.6 (released February 2026). Its safeguards allowed assistance with agentic cyber-exploit development, though CAISI noted GLM-5.2 appeared more resistant to jailbreaking and agent hijacking than other PRC open-weight models evaluated to that point — a distinction CAISI was explicit is largely moot once a model is self-hosted, since safeguards can be removed or bypassed outside the developer’s own serving infrastructure.
Kimi K3 (July 2026, joint with UK AISI): on ExploitBench, Kimi K3 scored 32% against GLM-5.2’s 24%, but achieved full automated compromise (ACE) on 0 of 41 sampled vulnerabilities, versus an average of 20 of 41 for the most cyber-capable models evaluated. On the TLO cyber range, Kimi K3 averaged 17 of 32 attack steps, against roughly 28.5 for the most capable U.S. models — though in one of ten attempts it completed the full 32-step range within the token budget. CAISI and UK AISI reported that Kimi K3’s safeguards did not prevent it from attempting cyber-exploit development or offensive operations when asked.
GLM-5.3 (September 2026): CAISI’s headline conclusion is that GLM-5.3 is “the most cyber-capable open-weight model released to date,” while remaining roughly four months behind U.S. frontier models in aggregate performance across the benchmark suite — 40.4% vs. 90.2% on SEC-Bench Pro, 61.1% vs. 100% on ExploitBench, 9.4% vs. 44.4% on ExploitGym, and 7.7% vs. 23.2% on OSS-Fuzz. Read across all three assessments, the trend line is a narrowing but still real gap: each successive open-weight release from a PRC lab has scored higher than the last, without yet closing on U.S. frontier capability.
Why open-weight, PRC-origin evaluation is a distinct governance question
Most of the AI-governance content in this cluster — safety frameworks, capability thresholds, pre-deployment testing agreements — concerns closed frontier labs that control their own model’s deployment: a lab can restrict API access, revoke a deployment, or update a safeguard centrally. Open-weight models remove that lever entirely. Once a lab publishes weights, as Z.ai and Moonshot AI both did for the models above, anyone can download and self-host the model, and any safeguard the developer shipped is only as durable as the effort required to strip it back out locally. CAISI’s own assessments make this point directly: a model’s safeguard behavior as tested through the developer’s own interface says little about what a self-hosted deployment can be made to do.
That is also why CAISI evaluates these releases on a recurring cadence rather than once: a governance framework built around reviewing a lab’s deployment decision has nothing to review here, so the assessment itself — repeated against each new release — is effectively the only checkpoint available for this category of model.
A NIKOLAI tie-in: naming what these assessments actually are
NIKOLAI, CASRAI’s own frontier-AI-safety dictionary, is an independent reference work — not affiliated with, endorsed by, or operated by NIST, CAISI, UK AISI, or any lab. It does not assess any system and issues no designation of its own. What it provides is disambiguated vocabulary for describing evaluation work like CAISI’s, and Track N5, “Evidence and evaluations” is built for exactly this: its elements let someone reading a CAISI report line up what happened against a common reference rather than relying on the report’s own one-off phrasing.
In NIKOLAI’s terms, each CAISI publication above is an evaluation — a structured assessment of a system against defined tasks — built from many individual evaluation runs, i.e. the ReAct-harness attempts against SEC-Bench Pro, ExploitBench, OSS-Fuzz, and TLO cyber-range tasks described above. CAISI’s choice to test closed comparator models with safeguards disabled, and to run open-weight models under maximum-reasoning settings inside an agent harness, is what NIKOLAI classifies as the assessment’s elicitation method — the conditions under which capability is drawn out of a model, which matters because two evaluations of the same model under different elicitation methods can produce genuinely different capability estimates, not just noise. This is CASRAI’s own reading applied to CAISI’s public methodology descriptions, not a crosswalk CAISI has reviewed or confirmed — under NIKOLAI’s Mapping Declarations system, every such row is a shadow mapping unless the organization itself has filed a declaration, and neither NIST nor CAISI has done so.
Frequently asked questions
Has CAISI assessed any open-weight models besides GLM-5.2, GLM-5.3, and Kimi K3?
Yes — CAISI’s publication record extends earlier in 2026 as well, including an evaluation of DeepSeek’s V4 Pro model. The three assessments covered in detail above were chosen because they fall inside a single four-month window and share closely comparable methodology, which is what makes the cadence itself visible.
Does a lower benchmark score mean an open-weight model is safe to self-host without concern?
CAISI’s own assessments don’t frame the results that way. Even the lowest-scoring model in this set (Kimi K3) achieved a nonzero success rate on the harder benchmarks and completed a full 32-step attack simulation in at least one attempt, and every assessment notes that self-hosting removes whatever safeguards the assessment observed through the developer’s own interface. The benchmark scores measure relative capability at a point in time, not an all-clear.
Is this the same thing as CAISI’s pre-deployment testing agreements with labs like OpenAI or Anthropic?
No. Pre-deployment testing agreements are bilateral arrangements CAISI has with individual labs, generally before a closed model ships. The open-weight assessments covered here happen after public release, are not the product of an agreement with the developer (Z.ai and Moonshot AI are not CAISI testing-agreement partners in the way U.S. labs are), and evaluate a model that’s already downloadable rather than one still under the developer’s deployment control.
Is NIKOLAI’s mapping of these assessments endorsed by CAISI or NIST?
No. NIKOLAI is CASRAI’s own, unendorsed reference work. Every crosswalk row connecting a NIKOLAI element to an organization’s actual practice is a shadow mapping — CASRAI’s independent reading of what that organization has published — unless the organization has filed an explicit Mapping Declaration confirming it. Neither NIST nor CAISI has done so for any element on this page.
Related reading
- Open-Weight AI Models: The Safety Case For vs. Against — the upstream policy debate this guide’s downstream monitoring work sits inside: whether release should happen at all, not just what CAISI finds once it has.







