Official figures position GLM-5.1's strengths in long-horizon code optimization and Agent tool loops, but scores depend heavily on harnesses such as OpenHands, Terminus, and Claude Code, as well as context management and specific parameters; they should not be treated directly as a ranking of bare models.
Suitable tasks: Engineering in real repositories, terminal tasks, code security, performance optimization, and long-running autonomous iteration.
Unsuitable tasks: Extrapolating long-horizon coding scores to general reasoning, mathematics, or another tool orchestration setup; the HLE/GPQA/AIME results do not lead across the board.
Applicable model version: GLM-5.1.
Applicable clients, Agents, or APIs: Z.ai API, OpenHands, Terminus-2, and Claude Code 2.1.x; the official release also provides local vLLM/SGLang paths.
Recommended reasoning tier and parameters: Fix parameters for the target benchmark; do not copy one harness's temperature/top_p/max_new_tokens to another client without rerunning the tests.
SWE-Bench Pro: OpenHands; temperature=1, top_p=0.95, max_new_tokens=32768, and 200K context.
NL2Repo: temperature=1.0, top_p=1.0, max_new_tokens=32768, and 200K context; rule-based prechecks and model judgment block malicious commands.
Terminal-Bench 2.0 Terminus-2: 3-hour timeout, temperature=1.0, top_p=1.0, max_new_tokens=8192, and 200K context; up to 16 CPUs and 32GB RAM.
Terminal-Bench Claude Code: Claude Code 2.1.69 think mode, temperature=1.0, top_p=0.95, and max_new_tokens=131072; with the wall-clock limit removed, scores are averaged over 5 runs.
CyberGym: Claude Code 2.1.56 think mode, without web tools; 1,507 tasks, single-run Pass@1, and a 250-minute timeout per task.
MCP-Atlas: 500-task public subset, think mode, a 10-minute timeout per task, and Gemini-3.0-Pro as the judge.
Long-horizon cases: VectorDBBench, KernelBench Level 3, and a Linux desktop web app; the latter uses a harness that has the model review its own output each round and continue improving it.
| Benchmark | GLM-5.1 | GLM-5 | Main comparisons |
|---|---|---|---|
| HLE | 31.0 | 30.5 | Claude Opus 4.6 36.7; GPT-5.4 39.8 |
| HLE w/ Tools | 52.3 | 50.4 | Claude 53.1*; GPT-5.4 52.1* |
| AIME 2026 | 95.3 | 95.4 | GPT-5.4 98.7 |
| GPQA-Diamond | 86.2 | 86.0 | Claude 91.3; GPT-5.4 92.0 |
| SWE-Bench Pro | 58.4 | 55.1 | GPT-5.4 57.7; Claude Opus 4.6 57.3; Gemini 54.2 |
| NL2Repo | 42.7 | 35.9 | Claude 49.8; GPT-5.4 41.3 |
| Terminal-Bench 2.0 Terminus-2 | 63.5 | 56.2 | Claude 65.4; Gemini 68.5 |
| CyberGym | 68.7 | 48.3 | Claude 66.6; GPT-5.4 66.3 |
| BrowseComp (without context management) | 68.0 | 62.0 | — |
| BrowseComp (with context management) | 79.3 | 75.9 | Claude 84.0; Gemini 85.9 |
| MCP-Atlas Public | 71.8 | 69.2 | Claude 73.8; GPT-5.4 67.2 |
Long-horizon cases: More than 600 optimization iterations and 6,000+ tool calls in VectorDBBench, from about 3,547 to 21,500 QPS, with Recall constrained to ≥95%; a final geometric mean of about 3.6× speedup on KernelBench Level 3, compared with about 4.2× for Claude Opus 4.6; and a Linux desktop web app that ran for 8 hours, continuously adding features and fixing interactions.
Compared with GLM-5, GLM-5.1's most credible gains appear in “working longer” and “still finding structural improvements in optimization loops with feedback.” SWE-Bench Pro 58.4, CyberGym 68.7, and VectorDBBench 21.5k QPS support trying it as an engineering Agent; the mathematics and general-reasoning data, however, show that it is not the overall frontier leader.
These are results self-reported by Z.ai and, unless otherwise noted, cannot be treated as independent controlled tests; the harness, judge, context trimming, and safety refusals all affect scores.
The 600-round/8-hour cases are individual showcase tasks; the full prompts, tool traces, failure samples, and costs have not all been disclosed.
Terminus-2, Claude Code, and best self-reported harness scores on Terminal-Bench are not interchangeable for comparison.
Some HLE/GPT/Claude comparisons carry asterisks; the official footnotes explicitly state that their sources or evaluation conditions differ. Do not treat the table as a set of fully homogeneous experiments.
Fix the model version, harness, tools, timeout, context, parameters, and resource limits, and save the configuration file.
First run public SWE-Bench Pro/NL2Repo/Terminal-Bench subsets, saving patches, tests, tool calls, and per-round metrics.
Use the same feedback loop for VectorDBBench/KernelBench, clearly specifying constraints, baselines, submission frequency, and stopping conditions.
Repeat code tasks multiple times at minimum; record tokens, cost, failures/fallbacks, and structural strategy changes, rather than only the final score.
Compare GLM-5, Opus/GPT, and other models with the same harness, reporting pure reasoning, coding, tool use, and long-horizon efficiency separately.
The official source publishes the benchmark table, parameters/resources for each task type, the three long-horizon scenarios—VectorDBBench, KernelBench, and Linux desktop—and figures including 58.4 SWE-Bench Pro, 63.5 Terminus-2, and 68.7 CyberGym.
The official source says GLM-5.1 “stays productive over longer sessions,” while also acknowledging that long-horizon optimization still faces local optima, coherence challenges across thousands of tool calls, and self-evaluation problems on tasks without metrics; this limits the scope of the conclusion.
GLM-5.1