NIST CAISI completed its assessment about three weeks after GLM-5.2 was released (on 2026-06-16, when Z.ai's predecessor, Zhipu AI, released the open weights), reaching independent conclusions from an official institutional perspective:
It was probably the most capable open-weight model when it was released ("probably the most capable open-weight AI model when it was released").
Its overall capabilities were comparable to those of GPT-5.2, released in December 2025.
Its cyber capabilities were comparable to those of Claude Opus 4.6, released in February 2026.
Its safeguards were mixed:
It allowed assistance with agentic cyber exploit development;
It blocked fewer sensitive biological queries than U.S. reference models;
But its robustness against agent hijacking and jailbreaking attacks may be higher than that of other evaluated PRC open-weight models;
Note: safeguards for open-weight models can all be circumvented when they are self-hosted.
The press release includes a graph comparing the overall capabilities of the strongest U.S. and Chinese models at release over time (Figure 1). On the y-axis, 400 points correspond to a 10-fold increase in the probability of solving a task; the methodology and details are in Appendices A1 and A4 of the complete assessment report. The body of the report was not published with the press release and must be obtained separately from the NIST CAISI site.
This assessment reflects an independent U.S. government perspective and is separate from Z.ai's official benchmark scores, providing cross-validation for the conclusion that "GLM-5.2 was the strongest open-weight model at release."
Evidence-level explanation: This is an independent institutional assessment with a complete methodology (although the body of the report was not published on the page collected this time); reproducing it requires obtaining CAISI's complete assessment report.
Use cases: It can be used to assess GLM-5.2's (1) position among open-weight models, (2) capability gap versus closed-source flagships (at the GPT-5.2 and Opus 4.6 level), (3) cyber capability risk level, and (4) safeguard strength—especially for compliance and security assessments of "whether to introduce this model" during model selection.
Scope: The conclusions are based on CAISI's own evaluation set, not a general-purpose ranking; the cyber-security capability conclusions concern sensitive uses, so keep the context in mind when citing them; the safeguard conclusions do not apply to self-hosted deployments (where they can be circumvented).
"GLM-5.2 was probably the most capable open-weight AI model when it was released."
"GLM-5.2's overall capabilities are similar to that of GPT-5.2, released in December 2025."
"GLM-5.2's cyber capabilities are similar to that of Opus 4.6, released in February 2026."
"GLM-5.2 appears potentially more robust against agent hijacking and jailbreaking attacks than other evaluated PRC open-… This is a necessary excerpt; read the original source for full context.
"Regardless of their robustness, safeguards for open-weight models can be circumvented when self-hosted."
GLM-5.2