Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGLM-5.1

GLM-5.1: Z.ai's Official Long-Horizon Engineering Benchmarks and Reproduction Conditions

Original source

Z.ai

AuthorZ.ai official

Source date2026-04-07

Tabbit curation2026-08-19

Read original

One-sentence takeaway

Official figures position GLM-5.1's strengths in long-horizon code optimization and Agent tool loops, but scores depend heavily on harnesses such as OpenHands, Terminus, and Claude Code, as well as context management and specific parameters; they should not be treated directly as a ranking of bare models.

Use cases

  • Suitable tasks: Engineering in real repositories, terminal tasks, code security, performance optimization, and long-running autonomous iteration.

  • Unsuitable tasks: Extrapolating long-horizon coding scores to general reasoning, mathematics, or another tool orchestration setup; the HLE/GPQA/AIME results do not lead across the board.

  • Applicable model version: GLM-5.1.

  • Applicable clients, Agents, or APIs: Z.ai API, OpenHands, Terminus-2, and Claude Code 2.1.x; the official release also provides local vLLM/SGLang paths.

  • Recommended reasoning tier and parameters: Fix parameters for the target benchmark; do not copy one harness's temperature/top_p/max_new_tokens to another client without rerunning the tests.

Test environment and input/configuration

  • SWE-Bench Pro: OpenHands; temperature=1, top_p=0.95, max_new_tokens=32768, and 200K context.

  • NL2Repo: temperature=1.0, top_p=1.0, max_new_tokens=32768, and 200K context; rule-based prechecks and model judgment block malicious commands.

  • Terminal-Bench 2.0 Terminus-2: 3-hour timeout, temperature=1.0, top_p=1.0, max_new_tokens=8192, and 200K context; up to 16 CPUs and 32GB RAM.

  • Terminal-Bench Claude Code: Claude Code 2.1.69 think mode, temperature=1.0, top_p=0.95, and max_new_tokens=131072; with the wall-clock limit removed, scores are averaged over 5 runs.

  • CyberGym: Claude Code 2.1.56 think mode, without web tools; 1,507 tasks, single-run Pass@1, and a 250-minute timeout per task.

  • MCP-Atlas: 500-task public subset, think mode, a 10-minute timeout per task, and Gemini-3.0-Pro as the judge.

  • Long-horizon cases: VectorDBBench, KernelBench Level 3, and a Linux desktop web app; the latter uses a harness that has the model review its own output each round and continue improving it.

Results

BenchmarkGLM-5.1GLM-5Main comparisons
HLE31.030.5Claude Opus 4.6 36.7; GPT-5.4 39.8
HLE w/ Tools52.350.4Claude 53.1*; GPT-5.4 52.1*
AIME 202695.395.4GPT-5.4 98.7
GPQA-Diamond86.286.0Claude 91.3; GPT-5.4 92.0
SWE-Bench Pro58.455.1GPT-5.4 57.7; Claude Opus 4.6 57.3; Gemini 54.2
NL2Repo42.735.9Claude 49.8; GPT-5.4 41.3
Terminal-Bench 2.0 Terminus-263.556.2Claude 65.4; Gemini 68.5
CyberGym68.748.3Claude 66.6; GPT-5.4 66.3
BrowseComp (without context management)68.062.0—
BrowseComp (with context management)79.375.9Claude 84.0; Gemini 85.9
MCP-Atlas Public71.869.2Claude 73.8; GPT-5.4 67.2

Long-horizon cases: More than 600 optimization iterations and 6,000+ tool calls in VectorDBBench, from about 3,547 to 21,500 QPS, with Recall constrained to ≥95%; a final geometric mean of about 3.6× speedup on KernelBench Level 3, compared with about 4.2× for Claude Opus 4.6; and a Linux desktop web app that ran for 8 hours, continuously adding features and fixing interactions.

Conclusion

Compared with GLM-5, GLM-5.1's most credible gains appear in “working longer” and “still finding structural improvements in optimization loops with feedback.” SWE-Bench Pro 58.4, CyberGym 68.7, and VectorDBBench 21.5k QPS support trying it as an engineering Agent; the mathematics and general-reasoning data, however, show that it is not the overall frontier leader.

Limitations

  • These are results self-reported by Z.ai and, unless otherwise noted, cannot be treated as independent controlled tests; the harness, judge, context trimming, and safety refusals all affect scores.

  • The 600-round/8-hour cases are individual showcase tasks; the full prompts, tool traces, failure samples, and costs have not all been disclosed.

  • Terminus-2, Claude Code, and best self-reported harness scores on Terminal-Bench are not interchangeable for comparison.

  • Some HLE/GPT/Claude comparisons carry asterisks; the official footnotes explicitly state that their sources or evaluation conditions differ. Do not treat the table as a set of fully homogeneous experiments.

Reproduction steps

  1. Fix the model version, harness, tools, timeout, context, parameters, and resource limits, and save the configuration file.

  2. First run public SWE-Bench Pro/NL2Repo/Terminal-Bench subsets, saving patches, tests, tool calls, and per-round metrics.

  3. Use the same feedback loop for VectorDBBench/KernelBench, clearly specifying constraints, baselines, submission frequency, and stopping conditions.

  4. Repeat code tasks multiple times at minimum; record tokens, cost, failures/fallbacks, and structural strategy changes, rather than only the final score.

  5. Compare GLM-5, Opus/GPT, and other models with the same harness, reporting pure reasoning, coding, tool use, and long-horizon efficiency separately.

Original evidence and data

The official source publishes the benchmark table, parameters/resources for each task type, the three long-horizon scenarios—VectorDBBench, KernelBench, and Linux desktop—and figures including 58.4 SWE-Bench Pro, 63.5 Terminus-2, and 68.7 CyberGym.

Source excerpt or observation (compliant short quote only)

The official source says GLM-5.1 “stays productive over longer sessions,” while also acknowledging that long-horizon optimization still faces local optima, coherence challenges across thousands of tool calls, and self-evaluation problems on tasks without metrics; this limits the scope of the conclusion.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GLM-5.1

Use and compare models in Tabbit

GLM-5.1

Related reviews

MediaSerenities AI2026-03-29

GLM-5.1: Serenities AI's Self-Reported Benchmarks and the Boundaries of Independent Validation

CommunityReddit r/LocalLLM

GLM-5.1: Reddit LocalLLM Real-World Coding and Context Experience

MediaArtificial Analysis2026-04-07

GLM-5.1: Artificial Analysis Independent Intelligence Index and Inference Throughput Benchmark

CommunityReddit r/opencodeCLI2026-05-15

GLM-5.1: OpenCode Three-Model Industrial Webpage Benchmark and Real-World Capability Boundaries

GLM-5.1

Related prompts

MediaZ.AI Developer Document / Z.ai2026-04-07

GLM-5.1: Long-horizon Agent and Claude Code Configuration