Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GLM-5.2 · Media / benchmark · Editorial analysis

NIST CAISI's Independent Capability Assessment of Z.ai GLM-5.2

NIST CAISI published its assessment on 2026-07-17 after completing it on 2026-07-08: GLM-5.2 was similar to GPT-5.2 overall and Opus 4.6 on cyber capability, while safeguards were mixed for agentic exploits and biological questions.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
Version GLM-5.2; NIST CAISI assessment completed 2026-07-08 and published 2026-07-17; institutional tests cover jailbreak robustness and safeguards, with full items/repeats in the report appendices.

Key data and applicable tasks

Core content summary

NIST CAISI completed its assessment about three weeks after GLM-5.2 was released (on 2026-06-16, when Z.ai's predecessor, Zhipu AI, released the open weights), reaching independent conclusions from an official institutional perspective:

Key conclusions (CAISI's own evaluation framework)

  1. It was probably the most capable open-weight model when it was released ("probably the most capable open-weight AI model when it was released").

  2. Its overall capabilities were comparable to those of GPT-5.2, released in December 2025.

  3. Its cyber capabilities were comparable to those of Claude Opus 4.6, released in February 2026.

  4. Its safeguards were mixed:

    • It allowed assistance with agentic cyber exploit development;

    • It blocked fewer sensitive biological queries than U.S. reference models;

    • But its robustness against agent hijacking and jailbreaking attacks may be higher than that of other evaluated PRC open-weight models;

    • Note: safeguards for open-weight models can all be circumvented when they are self-hosted.

Other information

  • The press release includes a graph comparing the overall capabilities of the strongest U.S. and Chinese models at release over time (Figure 1). On the y-axis, 400 points correspond to a 10-fold increase in the probability of solving a task; the methodology and details are in Appendices A1 and A4 of the complete assessment report. The body of the report was not published with the press release and must be obtained separately from the NIST CAISI site.

  • This assessment reflects an independent U.S. government perspective and is separate from Z.ai's official benchmark scores, providing cross-validation for the conclusion that "GLM-5.2 was the strongest open-weight model at release."

Evidence points and scope of applicability

  • Evidence-level explanation: This is an independent institutional assessment with a complete methodology (although the body of the report was not published on the page collected this time); reproducing it requires obtaining CAISI's complete assessment report.

  • Use cases: It can be used to assess GLM-5.2's (1) position among open-weight models, (2) capability gap versus closed-source flagships (at the GPT-5.2 and Opus 4.6 level), (3) cyber capability risk level, and (4) safeguard strength—especially for compliance and security assessments of "whether to introduce this model" during model selection.

  • Scope: The conclusions are based on CAISI's own evaluation set, not a general-purpose ranking; the cyber-security capability conclusions concern sensitive uses, so keep the context in mind when citing them; the safeguard conclusions do not apply to self-hosted deployments (where they can be circumvented).

Key quotations from the original

"GLM-5.2 was probably the most capable open-weight AI model when it was released."

"GLM-5.2's overall capabilities are similar to that of GPT-5.2, released in December 2025."

"GLM-5.2's cyber capabilities are similar to that of Opus 4.6, released in February 2026."

"GLM-5.2 appears potentially more robust against agent hijacking and jailbreaking attacks than other evaluated PRC open-… This is a necessary excerpt; read the original source for full context.

"Regardless of their robustness, safeguards for open-weight models can be circumvented when self-hosted."

What this supports

  • Supports the institution’s separate capability and safeguard findings.

What this does not support

  • Does not treat self-hosted open-weight safety as the same evaluated condition.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

NIST (National Institute of Standards and Technology) official news site · NIST CAISI (Center for AI Standards and Innovation) · Original publication date 2026-07-17 · Site edit date 2026-09-20

Open original source

GLM-5.2

Compare GLM-5.2 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GLM-5.2: What It Is, What It Costs, and Where It Fits

A sourced GLM-5.2 overview covering the June 2026 release, 1M context, open-weight deployment, API pricing boundaries, coding evidence and a safer pilot path.

Related reviews

Semgrep IDOR Benchmark: GLM-5.2 Results with a Prompt-Only Setup in Security Code AuditingSemgrep’s 2026-06-22 IDOR benchmark held dataset, evaluation, and prompt constant: GLM-5.2 reached 39% F1 in a Pydantic AI prompt-only harness at about $0.17 per vulnerability; this is not a general cyber score.GLM-5.2 Official Release Notes and Complete Benchmark Table (Z.ai Blog)Z.ai’s 2026-06-16 release positions GLM-5.2 as a 1M-context long-horizon flagship and reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-Bench Pro; it also discloses training-stage reward-hacking risk.Reddit Blind Code Review: GLM-5.2's Production-Readiness Score and Multi-Judge RecheckA Reddit VPS Manager blind review compared five models under one specification; Qwen 3.7 Plus first used a fixed 25-point rubric, followed by GPT Codex and Gemini 3.1 Pro rechecks; the sample is one project.Evening-Truth's Complaints About Z.AI Coding Plan Response Quality and Quantization SuspicionsEvening-Truth created a dedicated page in the prompt library to complain about the response quality of the Z.AI Coding Plan.GLM-5.2 Official Documentation: Overview and API Quick Start (docs.z.ai)The official standard integration configuration for GLM-5.2 is: model name `glm-5.2`, a 1M context window / 128K maximum output, `thinking.type: enabled` + `reasoning_effort: max`, and `temperature: 1.0`. You can copy the curl / Python examples directly to make your first call and review the typical use cases identified by the official documentation..GLM-5.2 Thinking Mode Configuration: Default Thinking / Interleaved Thinking / Preserved Thinking / Turn-level Thinking (Official)The official documentation states that thinking is enabled by default for GLM-5.2 (as with GLM-5.1/5/4.7), and provides four thinking modes: default thinking, interleaved thinking (thinking between tool calls), preserved thinking (retaining reasoning content across turns with `clear_thinking: false`), and turn-level thinking (an independent switch for each turn). It also highlights a key constraint for Agent integrations: historical `reasoning_content` must be returned unchanged..Official Configuration Guide for Migrating from GLM-5.1 / GLM-5 / GLM-4.x to GLM-5.2The official GLM-5.2 migration checklist and parameter configuration: change the model ID to `glm-5.2`; use the default `temperature` of 1.0 or default `top_p` of 0.95 (tune only one of the two); enable thinking by default; use `high` or `max` for `reasoning_effort`; configure streaming and streaming tool calls (`stream=true` + `tool_stream=true`) as specified by the official guidance; and use the included Python migration example directly..Using GLM-5.2 (zai-glm-5-2) Through Mistral: Third-Party Hosting Configuration and PricingMistral now hosts GLM-5.2 as a third-party open model (Public Preview, model ID `zai-glm-5-2`, 1M context / 128k output, with no modifications), so it can be accessed directly across the Mistral ecosystem (including Vibe CLI) using that ID, at $1.4 / $0.14 (cached input) / $4.4 (output) per million tokens..