Z.ai’s 2026-06-16 release positions GLM-5.2 as a 1M-context long-horizon flagship and reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-Bench Pro; it also discloses training-stage reward-hacking risk.
Z.ai official blog · Read evidenceGLM-5.2 · Reviews and evidence
Which GLM-5.2 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
NIST CAISI published its assessment on 2026-07-17 after completing it on 2026-07-08: GLM-5.2 was similar to GPT-5.2 overall and Opus 4.6 on cyber capability, while safeguards were mixed for agentic exploits and biological questions.
NIST (National Institute of Standards and Technology) official news site · Read evidenceSemgrep’s 2026-06-22 IDOR benchmark held dataset, evaluation, and prompt constant: GLM-5.2 reached 39% F1 in a Pydantic AI prompt-only harness at about $0.17 per vulnerability; this is not a general cyber score.
Semgrep official blog · Read evidenceFull reviews and related reading
Selected evidence
GLM-5.2 Official Release Notes and Complete Benchmark Table (Z.ai Blog)
Z.ai’s 2026-06-16 release positions GLM-5.2 as a 1M-context long-horizon flagship and reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-Bench Pro; it also discloses training-stage reward-hacking risk.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.2; official release 2026-06-16; tasks include Terminal-Bench 2.1, SWE-Bench Pro, and long-horizon engineering; full harness, repeats, and production success are undisclosed.
NIST CAISI's Independent Capability Assessment of Z.ai GLM-5.2
NIST CAISI published its assessment on 2026-07-17 after completing it on 2026-07-08: GLM-5.2 was similar to GPT-5.2 overall and Opus 4.6 on cyber capability, while safeguards were mixed for agentic exploits and biological questions.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.2; NIST CAISI assessment completed 2026-07-08 and published 2026-07-17; institutional tests cover jailbreak robustness and safeguards, with full items/repeats in the report appendices.
Semgrep IDOR Benchmark: GLM-5.2 Results with a Prompt-Only Setup in Security Code Auditing
Semgrep’s 2026-06-22 IDOR benchmark held dataset, evaluation, and prompt constant: GLM-5.2 reached 39% F1 in a Pydantic AI prompt-only harness at about $0.17 per vulnerability; this is not a general cyber score.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.2; Semgrep IDOR dataset of real open-source applications; same prompt/Pydantic AI harness; single configuration, 39% F1 and about $0.17/vulnerability, not a large rerun.
Reddit Blind Code Review: GLM-5.2's Production-Readiness Score and Multi-Judge Recheck
A Reddit VPS Manager blind review compared five models under one specification; Qwen 3.7 Plus first used a fixed 25-point rubric, followed by GPT Codex and Gemini 3.1 Pro rechecks; the sample is one project.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.2; one VPS Manager specification and five implementations; Qwen 3.7 Plus first scored a fixed rubric, then GPT Codex/Gemini 3.1 Pro rechecked; project count, prompts, and repeats are limited.
All sources
All sources
GLM-5.2 Official Release Notes and Complete Benchmark Table (Z.ai Blog)
Z.ai’s 2026-06-16 release positions GLM-5.2 as a 1M-context long-horizon flagship and reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-Bench Pro; it also discloses training-stage reward-hacking risk.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.2; official release 2026-06-16; tasks include Terminal-Bench 2.1, SWE-Bench Pro, and long-horizon engineering; full harness, repeats, and production success are undisclosed.
NIST CAISI's Independent Capability Assessment of Z.ai GLM-5.2
NIST CAISI published its assessment on 2026-07-17 after completing it on 2026-07-08: GLM-5.2 was similar to GPT-5.2 overall and Opus 4.6 on cyber capability, while safeguards were mixed for agentic exploits and biological questions.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.2; NIST CAISI assessment completed 2026-07-08 and published 2026-07-17; institutional tests cover jailbreak robustness and safeguards, with full items/repeats in the report appendices.
Semgrep IDOR Benchmark: GLM-5.2 Results with a Prompt-Only Setup in Security Code Auditing
Semgrep’s 2026-06-22 IDOR benchmark held dataset, evaluation, and prompt constant: GLM-5.2 reached 39% F1 in a Pydantic AI prompt-only harness at about $0.17 per vulnerability; this is not a general cyber score.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.2; Semgrep IDOR dataset of real open-source applications; same prompt/Pydantic AI harness; single configuration, 39% F1 and about $0.17/vulnerability, not a large rerun.
Reddit Blind Code Review: GLM-5.2's Production-Readiness Score and Multi-Judge Recheck
A Reddit VPS Manager blind review compared five models under one specification; Qwen 3.7 Plus first used a fixed 25-point rubric, followed by GPT Codex and Gemini 3.1 Pro rechecks; the sample is one project.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.2; one VPS Manager specification and five implementations; Qwen 3.7 Plus first scored a fixed rubric, then GPT Codex/Gemini 3.1 Pro rechecked; project count, prompts, and repeats are limited.
Independent Arena.ai Evaluation: GLM-5.2 (Max) Rankings in Code Arena / Agent Arena / Text Arena
Arena.ai evaluated GLM-5.2 (Max) in three types of in-platform evaluations in June 2026, with the following conclusions.
Unverified: the original source could not be rechecked.
- Model/version
- GLM-5.2; source title “Independent Arena.ai Evaluation: GLM-5.2 (Max) Rankings in Code Arena / Agent Arena / Text Arena”, with no cross-version merge.
- Task/harness
- Summary of key content Arena.ai evaluated GLM-5.2 (Max) in three types of in-platform evaluations in June 2026, with the following conclusions。 The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
X Field Test: GLM 5.2's Frontend Web Generation Rated "Far Ahead of the Current GPT Version" (@vista8)
Frontend developer @vista8 posted an X field test claiming GLM 5.2 produced visibly stronger frontend webpages than the current GPT version; the post includes a short side-by-side video and is a personal single-case observation, not a controlled benchmark.
Unverified: the original source could not be rechecked.
- Model/version
- GLM-5.2; source title “X Field Test: GLM 5.2's Frontend Web Generation Rated "Far Ahead of the Current GPT Version" (@vista8)”, with no cross-version merge.
- Task/harness
- Summary of key content Frontend developer @vista8 posted his field-test assessment on X: GLM 5.2 produces visibly better frontend webpages than the current GPT version; even a very good Skill cannot rescue GPT's "poor" f The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
r/LocalLLaMA Field Test: Running GLM 5.2 (744B MoE) Locally Without a GPU with the Colibrì Engine
u/LopsidedDot4557 shared a hands-on test of running GLM 5.2 (744B MoE) on a machine without a GPU, using colibrì, a pure-C engine that streams MoE experts from disk.
Unverified: the original source could not be rechecked.
- Model/version
- GLM-5.2; source title “r/LocalLLaMA Field Test: Running GLM 5.2 (744B MoE) Locally Without a GPU with the Colibrì Engine”, with no cross-version merge.
- Task/harness
- Core content summary u/LopsidedDot4557 shared a hands-on test of running GLM 5.2 (744B MoE) on a machine without a GPU, using colibrì, a pure-C engine that streams MoE experts from disk。 The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
r/PoeAI Discussion: Did GLM-5.2 Suddenly “Get Dumber”? — Community Evidence of Quality Differences Across Hosting Platforms
u/Friendly-Play-1953 posted in r/PoeAI that GLM-5.2 on Poe had consistently worked very well for creative writing, but over the past two or three days it had “lost half its IQ overnight”—even forgetting basic details across posts (the sa.
Unverified: the original source could not be rechecked.
- Model/version
- GLM-5.2; source title “r/PoeAI Discussion: Did GLM-5.2 Suddenly “Get Dumber”? — Community Evidence of Quality Differences Across Hosting Platforms”, with no cross-version merge.
- Task/harness
- Summary of key content u/Friendly-Play-1953 posted in r/PoeAI that GLM-5.2 on Poe had consistently worked very well for creative writing, but over the past two or three days it had “lost half its IQ overnight”—even forge The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
Evening-Truth's Complaints About Z.AI Coding Plan Response Quality and Quantization Suspicions
Evening-Truth created a dedicated page in the prompt library to complain about the response quality of the Z.AI Coding Plan.
Unverified: the original source could not be rechecked.
- Model/version
- GLM-5.2; source title “Evening-Truth's Complaints About Z.AI Coding Plan Response Quality and Quantization Suspicions”, with no cross-version merge.
- Task/harness
- Summary of key content Evening-Truth created a dedicated page in the prompt library to complain about the response quality of the Z.AI Coding Plan。 The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
r/openrouter Discussion: GLM-5.2 Free Endpoint (Hosted by Decart): Availability and Limitations
u/FreeTruck7609 discovered a free GLM endpoint hosted by Decart on OpenRouter (launched "about a day ago"). It had reliability problems at first, then stabilized at close to 100% availability..
Unverified: the original source could not be rechecked.
- Model/version
- GLM-5.2; source title “r/openrouter Discussion: GLM-5.2 Free Endpoint (Hosted by Decart): Availability and Limitations”, with no cross-version merge.
- Task/harness
- Summary of key content u/FreeTruck7609 discovered a free GLM endpoint hosted by Decart on OpenRouter (launched "about a day ago"). It had reliability problems at first, then stabilized at close to 100% availability.。 The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
Hugging Face Security Incident Forensics: GLM-5.2 Used for Self-Hosted Attack Log Analysis (Real-World Project Report)
Hugging Face disclosed that it suffered an intrusion in July 2026 initiated by an autonomous Agent framework (the adversary operated as an "agentic attacker," executing thousands of automated actions in a large number of short-lived sandbo.
Unverified: the original source could not be rechecked.
- Model/version
- GLM-5.2; source title “Hugging Face Security Incident Forensics: GLM-5.2 Used for Self-Hosted Attack Log Analysis (Real-World Project Report)”, with no cross-version merge.
- Task/harness
- Core content summary Hugging Face disclosed that it suffered an intrusion in July 2026 initiated by an autonomous Agent framework (the adversary operated as an "agentic attacker," executing thousands of automated actions The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
GLM-5.2
Compare GLM-5.2 in Tabbit
Model access, features, and permissions depend on your current client account.