Skip to main content
v2026.11,858 entries · CC-BY 4.0

CAISI’s Evaluation Cadence for Open-Weight and PRC-Origin AI Models

CAISI has published cyber-capability assessments of three PRC-lab open-weight models in four months — GLM-5.2, Kimi K3 (with UK AISI), and GLM-5.3 — a real, ongoing evaluation cadence, not a one-off report. What the methodology and results actually show.

Written and maintained by CASRAI Editorial Board

Last updated

On September 17, 2026, NIST’s Center for AI Standards and Innovation (CAISI) published its assessment of Z.ai’s GLM-5.3 — the model Z.ai had released just over a month earlier, on August 14, 2026. It was not CAISI’s first look at a Chinese-lab open-weight model, and it will not be its last: over the four months before it, CAISI (sometimes jointly with the UK AI Security Institute, UK AISI) had already published a cyber-capability assessment of Z.ai’s GLM-5.2 (July 17, 2026) and a preliminary joint assessment of Moonshot AI’s Kimi K3 (July 23, 2026). That is a real, dated, primary-sourced cadence — not a single reactive report — and it answers a governance question that is genuinely distinct from the closed-frontier-lab safety-framework content that dominates most AI-governance coverage: what happens when a highly capable model’s weights are simply published, so there is no API to gate and no deployment decision to review?

This guide covers what CAISI has actually assessed so far, how its evaluation methodology works across those assessments, and why open-weight, PRC-origin model risk is its own governance category. For CAISI as an institution — its mandate, its 2025 rename, and its relationship to the UK AI Security Institute — see our companion guide. For the mechanics of CAISI’s pre-deployment testing agreements with individual labs, see our guide to those agreements.

The cadence so far

Three assessments, each targeting a newly released open-weight model from a PRC-based lab, published in close succession:

Published Model Developer Model released Collaboration
July 17, 2026 GLM-5.2 Z.ai (formerly Zhipu AI) June 16, 2026 CAISI
July 23, 2026 Kimi K3 Moonshot AI July 16, 2026 UK AISI & CAISI (preliminary joint assessment, updated August 28, 2026)
September 17, 2026 GLM-5.3 Z.ai August 14, 2026 CAISI

In each case, CAISI (and UK AISI, for the joint Kimi K3 assessment) moved from a model’s public release to a published cyber-capability assessment in roughly four to five weeks. That turnaround, repeated three times against three different labs’ releases inside one quarter, is the actual evidence of an ongoing program rather than a one-off response to a single model.

How CAISI evaluates open-weight cyber capability

Across these assessments, CAISI’s approach centers on aggregating performance across many individual tasks into a single capability estimate, rather than reporting raw pass rates in isolation. The GLM-5.3 and Kimi K3 assessments both describe an approach built on (or “inspired by”) Item Response Theory (IRT), producing a cyber capability index in which a fixed-size increase in score corresponds to a fixed multiplicative increase in the odds of solving a task — a scale designed to let capability be compared across models and over time, not just within one report.

The specific tasks vary by assessment but cluster around a few benchmark families:

  • ExploitBench — a public benchmark (Carnegie Mellon University) measuring how far a model can progress along the software-exploitation ladder against a fixed set of post-2023 V8 JavaScript-engine vulnerabilities.
  • SEC-Bench Pro and CAISI OSS-Fuzz — locating and exploiting vulnerabilities in browser engines (V8, SpiderMonkey) and in widely used open-source projects, including a large set of non-public tasks CAISI maintains specifically so models cannot have trained on the answers.
  • ExploitGym Userspace — developing exploits from bugs in open-source projects at larger scale.
  • “The Last Ones” (TLO) cyber range — a simulated 32-step corporate-network attack path across roughly 20 hosts and 4 subnets, used in the Kimi K3 assessment to measure sustained, multi-stage offensive capability rather than single-exploit development.

Models are run inside an agent harness (a ReAct-style loop with bash and Python tools, turn limits in the low hundreds, maximum reasoning settings enabled) rather than queried for one-shot answers — the assessments are explicitly measuring what an agentic system built on the model can accomplish, not what the base model can describe. For the closed U.S. models used as comparators, CAISI evaluates with system-level safeguards disabled, on the stated basis that this is necessary to measure maximal capability rather than a lab’s current refusal behavior; the open-weight models under assessment are, by construction, already runnable without a vendor’s safeguards once self-hosted.

What the three assessments found

GLM-5.2 (July 2026): CAISI’s assessment describes GLM-5.2 as “probably the most capable open-weight AI model when it was released,” with overall capability comparable to GPT-5.2 (released December 2025) and cyber capability comparable to Claude Opus 4.6 (released February 2026). Its safeguards allowed assistance with agentic cyber-exploit development, though CAISI noted GLM-5.2 appeared more resistant to jailbreaking and agent hijacking than other PRC open-weight models evaluated to that point — a distinction CAISI was explicit is largely moot once a model is self-hosted, since safeguards can be removed or bypassed outside the developer’s own serving infrastructure.

Kimi K3 (July 2026, joint with UK AISI): on ExploitBench, Kimi K3 scored 32% against GLM-5.2’s 24%, but achieved full automated compromise (ACE) on 0 of 41 sampled vulnerabilities, versus an average of 20 of 41 for the most cyber-capable models evaluated. On the TLO cyber range, Kimi K3 averaged 17 of 32 attack steps, against roughly 28.5 for the most capable U.S. models — though in one of ten attempts it completed the full 32-step range within the token budget. CAISI and UK AISI reported that Kimi K3’s safeguards did not prevent it from attempting cyber-exploit development or offensive operations when asked.

GLM-5.3 (September 2026): CAISI’s headline conclusion is that GLM-5.3 is “the most cyber-capable open-weight model released to date,” while remaining roughly four months behind U.S. frontier models in aggregate performance across the benchmark suite — 40.4% vs. 90.2% on SEC-Bench Pro, 61.1% vs. 100% on ExploitBench, 9.4% vs. 44.4% on ExploitGym, and 7.7% vs. 23.2% on OSS-Fuzz. Read across all three assessments, the trend line is a narrowing but still real gap: each successive open-weight release from a PRC lab has scored higher than the last, without yet closing on U.S. frontier capability.

Why open-weight, PRC-origin evaluation is a distinct governance question

Most of the AI-governance content in this cluster — safety frameworks, capability thresholds, pre-deployment testing agreements — concerns closed frontier labs that control their own model’s deployment: a lab can restrict API access, revoke a deployment, or update a safeguard centrally. Open-weight models remove that lever entirely. Once a lab publishes weights, as Z.ai and Moonshot AI both did for the models above, anyone can download and self-host the model, and any safeguard the developer shipped is only as durable as the effort required to strip it back out locally. CAISI’s own assessments make this point directly: a model’s safeguard behavior as tested through the developer’s own interface says little about what a self-hosted deployment can be made to do.

That is also why CAISI evaluates these releases on a recurring cadence rather than once: a governance framework built around reviewing a lab’s deployment decision has nothing to review here, so the assessment itself — repeated against each new release — is effectively the only checkpoint available for this category of model.

A NIKOLAI tie-in: naming what these assessments actually are

NIKOLAI, CASRAI’s own frontier-AI-safety dictionary, is an independent reference work — not affiliated with, endorsed by, or operated by NIST, CAISI, UK AISI, or any lab. It does not assess any system and issues no designation of its own. What it provides is disambiguated vocabulary for describing evaluation work like CAISI’s, and Track N5, “Evidence and evaluations” is built for exactly this: its elements let someone reading a CAISI report line up what happened against a common reference rather than relying on the report’s own one-off phrasing.

In NIKOLAI’s terms, each CAISI publication above is an evaluation — a structured assessment of a system against defined tasks — built from many individual evaluation runs, i.e. the ReAct-harness attempts against SEC-Bench Pro, ExploitBench, OSS-Fuzz, and TLO cyber-range tasks described above. CAISI’s choice to test closed comparator models with safeguards disabled, and to run open-weight models under maximum-reasoning settings inside an agent harness, is what NIKOLAI classifies as the assessment’s elicitation method — the conditions under which capability is drawn out of a model, which matters because two evaluations of the same model under different elicitation methods can produce genuinely different capability estimates, not just noise. This is CASRAI’s own reading applied to CAISI’s public methodology descriptions, not a crosswalk CAISI has reviewed or confirmed — under NIKOLAI’s Mapping Declarations system, every such row is a shadow mapping unless the organization itself has filed a declaration, and neither NIST nor CAISI has done so.

Frequently asked questions

Has CAISI assessed any open-weight models besides GLM-5.2, GLM-5.3, and Kimi K3?

Yes — CAISI’s publication record extends earlier in 2026 as well, including an evaluation of DeepSeek’s V4 Pro model. The three assessments covered in detail above were chosen because they fall inside a single four-month window and share closely comparable methodology, which is what makes the cadence itself visible.

Does a lower benchmark score mean an open-weight model is safe to self-host without concern?

CAISI’s own assessments don’t frame the results that way. Even the lowest-scoring model in this set (Kimi K3) achieved a nonzero success rate on the harder benchmarks and completed a full 32-step attack simulation in at least one attempt, and every assessment notes that self-hosting removes whatever safeguards the assessment observed through the developer’s own interface. The benchmark scores measure relative capability at a point in time, not an all-clear.

Is this the same thing as CAISI’s pre-deployment testing agreements with labs like OpenAI or Anthropic?

No. Pre-deployment testing agreements are bilateral arrangements CAISI has with individual labs, generally before a closed model ships. The open-weight assessments covered here happen after public release, are not the product of an agreement with the developer (Z.ai and Moonshot AI are not CAISI testing-agreement partners in the way U.S. labs are), and evaluate a model that’s already downloadable rather than one still under the developer’s deployment control.

Is NIKOLAI’s mapping of these assessments endorsed by CAISI or NIST?

No. NIKOLAI is CASRAI’s own, unendorsed reference work. Every crosswalk row connecting a NIKOLAI element to an organization’s actual practice is a shadow mapping — CASRAI’s independent reading of what that organization has published — unless the organization has filed an explicit Mapping Declaration confirming it. Neither NIST nor CAISI has done so for any element on this page.

Related reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about CAISI’s Evaluation Cadence for Open-Weight and PRC-Origin AI Models

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →