Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GLM-5.2 · Official source · Vendor report

GLM-5.2 Official Release Notes and Complete Benchmark Table (Z.ai Blog)

Z.ai’s 2026-06-16 release positions GLM-5.2 as a 1M-context long-horizon flagship and reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-Bench Pro; it also discloses training-stage reward-hacking risk.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Conditions
Version GLM-5.2; official release 2026-06-16; tasks include Terminal-Bench 2.1, SWE-Bench Pro, and long-horizon engineering; full harness, repeats, and production success are undisclosed.

Key data and applicable tasks

Core content summary

GLM-5.2 is Z.ai's flagship model for "long-horizon tasks." Its key selling points are a genuinely usable 1M-token context, flexible reasoning modes, and the MIT open-source license (with no regional restrictions). According to the official account, it delivers a substantial improvement over GLM-5.1 on long-horizon tasks and is the first model to "reliably withstand engineering pressure" with a 1M context.

Release highlights

  • 1M context: Training on a 1M context was substantially expanded for coding-agent scenarios (large-scale implementation, automated research, performance optimization, and complex debugging), emphasizing that "a 1M context must be engineering-usable, not merely able to accept more tokens."

  • IndexShare architecture: One indexer is shared by every four layers of sparse attention, reducing FLOPs per token by 2.9× at a 1M context length; improvements to the MTP (multi-token prediction) layer are used for speculative decoding, increasing acceptance length by up to 20% (ablation: baseline 4.56 → +IndexShare+KV Share 5.10 → +Rejection Sampling 5.29 → +End-to-end TV Loss 5.47).

  • Inference serving: LayerSplit provides fine-grained memory management and parallelism; long-context kernels and CPU-side cache management are optimized, with throughput advantages growing as the context gets longer.

  • Agentic RL (slime framework): Parallel OPD training combines 10+ expert models, with the entire OPD training taking about two days; critic-based PPO (single rollout, token-level advantage) is introduced to support compaction; coding RL adds two-stage anti-cheating (rule-based filter + LLM judge), intercepting hack calls online and returning fake data so the rollout can continue.

  • Access: GLM Coding Plan (model name GLM-5.2; use GLM-5.2[1m] in Claude Code to enable a 1M context); Z.ai chat; open-source weights on HuggingFace / ModelScope (supporting transformers, vLLM, SGLang, xLLM, and ktransformers).

Complete benchmark table (vendor-reported, 2026-06-16)

BenchmarkGLM-5.2GLM-5.1Qwen3.7-MaxMiniMax M3DeepSeek-V4-ProClaude Opus 4.8GPT-5.5Gemini 3.1 Pro
HLE40.531.041.437.037.749.8*41.4*45.0
HLE w/ Tools54.752.353.5-48.257.9*52.2*51.4*
CritPt20.94.613.43.712.920.927.117.7
AIME 202699.295.397.0-94.695.798.398.2
HMMT Nov. 202594.494.095.084.494.496.596.594.8
HMMT Feb. 202692.582.697.184.495.296.796.787.3
IMOAnswerBench91.083.890.0-89.883.5-81.0
GPQA-Diamond91.286.290.093.090.193.693.694.3
SWE-bench Pro62.158.460.659.055.469.258.654.2
NL2Repo48.942.747.242.135.569.750.733.4
DeepSWE46.218.018.020.08.058.070.010.0
ProgramBench63.750.9--47.871.970.839.5
Terminal-Bench 2.1 (Terminus-2)81.063.575.065.064.085.084.074.0
Terminal-Bench 2.1 (Best Harness)82.7 (Claude Code)69 (Claude Code)---78.9 (Claude Code)83.4 (Codex)70.7 (Gemini CLI)
FrontierSWE (Dominance, 2026/06/16)74.430.5--29.075.172.639.6
PostTrainBench34.320.1---37.228.421.6
SWE-Marathon13.01.0---26.012.04.0
MCP-Atlas (Public Set)76.871.876.474.273.677.875.369.2
Tool-Decathlon48.240.7--52.859.955.648.8

(* Full-set scores.)

Three long-horizon benchmarks (official figures)

  • FrontierSWE (Proximal evaluation, 1M context + Max mode + 128K output): 74.4, only about 1% behind Opus 4.8 (75.1), about 1% ahead of GPT-5.5 (72.6), and about 11% ahead of Opus 4.7.

  • PostTrainBench (each agent gets an H100, measuring how much it can improve a small model through post-training): 34.3, second only to Opus 4.8 (37.2).

  • SWE-Marathon (ultra-long-horizon tasks: writing a compiler, optimizing kernels, and developing production-grade services): 13.0, about 13% behind Opus 4.8 (26.0) (as stated in the original), and still the top open-source model.

Official evaluation-method footnotes (key to reproducibility)

  • HLE and other reasoning tasks: temperature=1.0, top_p=0.95, maximum generation length 163,840; the text-only subset is reported by default.

  • SWE-bench Pro: OpenHands + custom instruction prompt, temperature=1, top_p=1, max_new_tokens=32k, 400K context.

  • DeepSWE: official pier evaluation framework + mini-swe-agent, temperature=1.0, top_p=1.0, timeout=2h, 400K context, isolated container (2 CPUs, 8GB RAM, no network).

  • Terminal-Bench 2.1: Terminus-2 framework (parser=json, timeout=4h, max_new_tokens=48k, max_episodes=500, 256K context, 4 CPU / 8GB RAM limit); the Claude Code version uses a transparent proxy to raise max_new_tokens to 128k, removes the wall-clock limit, and averages five runs.

  • MCP-Atlas: think mode, a public subset of 500 tasks, a 10-minute timeout per task, and Gemini-3.0-Pro as the judge.

  • Usage note: Different models may use different harnesses, so cross-row comparisons should be made cautiously.

Quotas and costs (Coding Plan figures)

  • Quota consumption is 3× during peak periods and 2× off-peak; under the limited-time promotion (through the end of September), off-peak usage is counted as 1×. Peak hours are 14:00–18:00 (UTC+8) every day.

  • The official recommendation for coding tasks is Max mode.

Key quotations from the original

"Supporting long-horizon tasks starts with making long context engineering-usable: the model must maintain quality acros… This is a necessary excerpt; read the original source for full context.

"A 1M context is easy to claim, but much harder to keep reliable under real engineering pressure."

"Across all three benchmarks, GLM-5.2 is the highest-ranked open-source model, showing that its 1M context has translate… This is a necessary excerpt; read the original source for full context.

Scope and limitations

  • All benchmark scores are vendor-reported, including results run by third-party evaluation organizations (Proximal / PostTrainBench / Abundant AI); when comparing with closed-source models, note that the models may use different harnesses, contexts, and modes.

  • When citing external scores, it is recommended to provide the evaluation configurations from the official footnotes as well; the 5.2 baselines published officially when GLM-5.3 was released (such as Terminal-Bench 3.0 4.6 and DeepSWE v1.1 46.2) use different benchmarks from the 2.x versions in this article, so distinguish the versions carefully.

What this supports

  • Supports the official context, benchmark, and reward-hacking disclosure.

What this does not support

  • Does not make official scores stable across providers or a safety guarantee.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Z.ai official blog · Z.ai Research · Original publication date 2026-06-16 · Site edit date 2026-09-20

Open original source

GLM-5.2

Compare GLM-5.2 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GLM-5.2: What It Is, What It Costs, and Where It Fits

A sourced GLM-5.2 overview covering the June 2026 release, 1M context, open-weight deployment, API pricing boundaries, coding evidence and a safer pilot path.

Related reviews

NIST CAISI's Independent Capability Assessment of Z.ai GLM-5.2NIST CAISI published its assessment on 2026-07-17 after completing it on 2026-07-08: GLM-5.2 was similar to GPT-5.2 overall and Opus 4.6 on cyber capability, while safeguards were mixed for agentic exploits and biological questions.Semgrep IDOR Benchmark: GLM-5.2 Results with a Prompt-Only Setup in Security Code AuditingSemgrep’s 2026-06-22 IDOR benchmark held dataset, evaluation, and prompt constant: GLM-5.2 reached 39% F1 in a Pydantic AI prompt-only harness at about $0.17 per vulnerability; this is not a general cyber score.Reddit Blind Code Review: GLM-5.2's Production-Readiness Score and Multi-Judge RecheckA Reddit VPS Manager blind review compared five models under one specification; Qwen 3.7 Plus first used a fixed 25-point rubric, followed by GPT Codex and Gemini 3.1 Pro rechecks; the sample is one project.Independent Arena.ai Evaluation: GLM-5.2 (Max) Rankings in Code Arena / Agent Arena / Text ArenaArena.ai evaluated GLM-5.2 (Max) in three types of in-platform evaluations in June 2026, with the following conclusions.GLM-5.2 Official Documentation: Overview and API Quick Start (docs.z.ai)The official standard integration configuration for GLM-5.2 is: model name `glm-5.2`, a 1M context window / 128K maximum output, `thinking.type: enabled` + `reasoning_effort: max`, and `temperature: 1.0`. You can copy the curl / Python examples directly to make your first call and review the typical use cases identified by the official documentation..GLM-5.2 Thinking Mode Configuration: Default Thinking / Interleaved Thinking / Preserved Thinking / Turn-level Thinking (Official)The official documentation states that thinking is enabled by default for GLM-5.2 (as with GLM-5.1/5/4.7), and provides four thinking modes: default thinking, interleaved thinking (thinking between tool calls), preserved thinking (retaining reasoning content across turns with `clear_thinking: false`), and turn-level thinking (an independent switch for each turn). It also highlights a key constraint for Agent integrations: historical `reasoning_content` must be returned unchanged..Official Configuration Guide for Migrating from GLM-5.1 / GLM-5 / GLM-4.x to GLM-5.2The official GLM-5.2 migration checklist and parameter configuration: change the model ID to `glm-5.2`; use the default `temperature` of 1.0 or default `top_p` of 0.95 (tune only one of the two); enable thinking by default; use `high` or `max` for `reasoning_effort`; configure streaming and streaming tool calls (`stream=true` + `tool_stream=true`) as specified by the official guidance; and use the included Python migration example directly..Using GLM-5.2 (zai-glm-5-2) Through Mistral: Third-Party Hosting Configuration and PricingMistral now hosts GLM-5.2 as a third-party open model (Public Preview, model ID `zai-glm-5-2`, 1M context / 128k output, with no modifications), so it can be accessed directly across the Mistral ecosystem (including Vibe CLI) using that ID, at $1.4 / $0.14 (cached input) / $4.4 (output) per million tokens..