GLM-5.1 · Official source · Vendor report
Z.AI’s 2026-04-07 material claims up to 8 hours of sustained execution, 58.4 on SWE-Bench Pro, and 3.6× geometric-mean speedup on KernelBench Level 3; results depend on OpenHands/Terminus/Claude Code harnesses.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
Official figures position GLM-5.1's strengths in long-horizon code optimization and Agent tool loops, but scores depend heavily on harnesses such as OpenHands, Terminus, and Claude Code, as well as context management and specific parameters; they should not be treated directly as a ranking of bare models.
Suitable tasks: Engineering in real repositories, terminal tasks, code security, performance optimization, and long-running autonomous iteration.
Unsuitable tasks: Extrapolating long-horizon coding scores to general reasoning, mathematics, or another tool orchestration setup; the HLE/GPQA/AIME results do not lead across the board.
Applicable model version: GLM-5.1.
Applicable clients, Agents, or APIs: Z.ai API, OpenHands, Terminus-2, and Claude Code 2.1.x; the official release also provides local vLLM/SGLang paths.
Recommended reasoning tier and parameters: Fix parameters for the target benchmark; do not copy one harness's temperature/top_p/max_new_tokens to another client without rerunning the tests.
SWE-Bench Pro: OpenHands; temperature=1, top_p=0.95, max_new_tokens=32768, and 200K context.
NL2Repo: temperature=1.0, top_p=1.0, max_new_tokens=32768, and 200K context; rule-based prechecks and model judgment block malicious commands.
Terminal-Bench 2.0 Terminus-2: 3-hour timeout, temperature=1.0, top_p=1.0, max_new_tokens=8192, and 200K context; up to 16 CPUs and 32GB RAM.
Terminal-Bench Claude Code: Claude Code 2.1.69 think mode, temperature=1.0, top_p=0.95, and max_new_tokens=131072; with the wall-clock limit removed, scores are averaged over 5 runs.
CyberGym: Claude Code 2.1.56 think mode, without web tools; 1,507 tasks, single-run Pass@1, and a 250-minute timeout per task.
MCP-Atlas: 500-task public subset, think mode, a 10-minute timeout per task, and Gemini-3.0-Pro as the judge.
Long-horizon cases: VectorDBBench, KernelBench Level 3, and a Linux desktop web app; the latter uses a harness that has the model review its own output each round and continue improving it.
| Benchmark | GLM-5.1 | GLM-5 | Main comparisons |
|---|---|---|---|
| HLE | 31.0 | 30.5 | Claude Opus 4.6 36.7; GPT-5.4 39.8 |
| HLE w/ Tools | 52.3 | 50.4 | Claude 53.1*; GPT-5.4 52.1* |
| AIME 2026 | 95.3 | 95.4 | GPT-5.4 98.7 |
| GPQA-Diamond | 86.2 | 86.0 | Claude 91.3; GPT-5.4 92.0 |
| SWE-Bench Pro | 58.4 | 55.1 | GPT-5.4 57.7; Claude Opus 4.6 57.3; Gemini 54.2 |
| NL2Repo | 42.7 | 35.9 | Claude 49.8; GPT-5.4 41.3 |
| Terminal-Bench 2.0 Terminus-2 | 63.5 | 56.2 | Claude 65.4; Gemini 68.5 |
| CyberGym | 68.7 | 48.3 | Claude 66.6; GPT-5.4 66.3 |
| BrowseComp (without context management) | 68.0 | 62.0 | — |
| BrowseComp (with context management) | 79.3 | 75.9 | Claude 84.0; Gemini 85.9 |
| MCP-Atlas Public | 71.8 | 69.2 | Claude 73.8; GPT-5.4 67.2 |
Long-horizon cases: More than 600 optimization iterations and 6,000+ tool calls in VectorDBBench, from about 3,547 to 21,500 QPS, with Recall constrained to ≥95%; a final geometric mean of about 3.6× speedup on KernelBench Level 3, compared with about 4.2× for Claude Opus 4.6; and a Linux desktop web app that ran for 8 hours, continuously adding features and fixing interactions.
Compared with GLM-5, GLM-5.1's most credible gains appear in “working longer” and “still finding structural improvements in optimization loops with feedback.” SWE-Bench Pro 58.4, CyberGym 68.7, and VectorDBBench 21.5k QPS support trying it as an engineering Agent; the mathematics and general-reasoning data, however, show that it is not the overall frontier leader.
These are results self-reported by Z.ai and, unless otherwise noted, cannot be treated as independent controlled tests; the harness, judge, context trimming, and safety refusals all affect scores.
The 600-round/8-hour cases are individual showcase tasks; the full prompts, tool traces, failure samples, and costs have not all been disclosed.
Terminus-2, Claude Code, and best self-reported harness scores on Terminal-Bench are not interchangeable for comparison.
Some HLE/GPT/Claude comparisons carry asterisks; the official footnotes explicitly state that their sources or evaluation conditions differ. Do not treat the table as a set of fully homogeneous experiments.
Fix the model version, harness, tools, timeout, context, parameters, and resource limits, and save the configuration file.
First run public SWE-Bench Pro/NL2Repo/Terminal-Bench subsets, saving patches, tests, tool calls, and per-round metrics.
Use the same feedback loop for VectorDBBench/KernelBench, clearly specifying constraints, baselines, submission frequency, and stopping conditions.
Repeat code tasks multiple times at minimum; record tokens, cost, failures/fallbacks, and structural strategy changes, rather than only the final score.
Compare GLM-5, Opus/GPT, and other models with the same harness, reporting pure reasoning, coding, tool use, and long-horizon efficiency separately.
The official source publishes the benchmark table, parameters/resources for each task type, the three long-horizon scenarios—VectorDBBench, KernelBench, and Linux desktop—and figures including 58.4 SWE-Bench Pro, 63.5 Terminus-2, and 68.7 CyberGym.
The official source says GLM-5.1 “stays productive over longer sessions,” while also acknowledging that long-horizon optimization still faces local optima, coherence challenges across thousands of tool calls, and self-evaluation problems on tasks without metrics; this limits the scope of the conclusion.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Z.ai · Z.ai official · Original publication date 2026-04-07 · Site edit date 2026-09-20
Open original sourceGLM-5.1
Download the Tabbit client to check model access