Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GLM-5.1 · Official source · Vendor report

GLM-5.1: Z.ai's Official Long-Horizon Engineering Benchmarks and Reproduction Conditions

Z.AI’s 2026-04-07 material claims up to 8 hours of sustained execution, 58.4 on SWE-Bench Pro, and 3.6× geometric-mean speedup on KernelBench Level 3; results depend on OpenHands/Terminus/Claude Code harnesses.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Conditions
Version GLM-5.1; official release 2026-04-07; tasks include SWE-Bench Pro, KernelBench L3, and an 8-hour engineering loop; harness, context management, and parameters control comparability.

Key data and applicable tasks

One-sentence takeaway

Official figures position GLM-5.1's strengths in long-horizon code optimization and Agent tool loops, but scores depend heavily on harnesses such as OpenHands, Terminus, and Claude Code, as well as context management and specific parameters; they should not be treated directly as a ranking of bare models.

Use cases

  • Suitable tasks: Engineering in real repositories, terminal tasks, code security, performance optimization, and long-running autonomous iteration.

  • Unsuitable tasks: Extrapolating long-horizon coding scores to general reasoning, mathematics, or another tool orchestration setup; the HLE/GPQA/AIME results do not lead across the board.

  • Applicable model version: GLM-5.1.

  • Applicable clients, Agents, or APIs: Z.ai API, OpenHands, Terminus-2, and Claude Code 2.1.x; the official release also provides local vLLM/SGLang paths.

  • Recommended reasoning tier and parameters: Fix parameters for the target benchmark; do not copy one harness's temperature/top_p/max_new_tokens to another client without rerunning the tests.

Test environment and input/configuration

  • SWE-Bench Pro: OpenHands; temperature=1, top_p=0.95, max_new_tokens=32768, and 200K context.

  • NL2Repo: temperature=1.0, top_p=1.0, max_new_tokens=32768, and 200K context; rule-based prechecks and model judgment block malicious commands.

  • Terminal-Bench 2.0 Terminus-2: 3-hour timeout, temperature=1.0, top_p=1.0, max_new_tokens=8192, and 200K context; up to 16 CPUs and 32GB RAM.

  • Terminal-Bench Claude Code: Claude Code 2.1.69 think mode, temperature=1.0, top_p=0.95, and max_new_tokens=131072; with the wall-clock limit removed, scores are averaged over 5 runs.

  • CyberGym: Claude Code 2.1.56 think mode, without web tools; 1,507 tasks, single-run Pass@1, and a 250-minute timeout per task.

  • MCP-Atlas: 500-task public subset, think mode, a 10-minute timeout per task, and Gemini-3.0-Pro as the judge.

  • Long-horizon cases: VectorDBBench, KernelBench Level 3, and a Linux desktop web app; the latter uses a harness that has the model review its own output each round and continue improving it.

Results

BenchmarkGLM-5.1GLM-5Main comparisons
HLE31.030.5Claude Opus 4.6 36.7; GPT-5.4 39.8
HLE w/ Tools52.350.4Claude 53.1*; GPT-5.4 52.1*
AIME 202695.395.4GPT-5.4 98.7
GPQA-Diamond86.286.0Claude 91.3; GPT-5.4 92.0
SWE-Bench Pro58.455.1GPT-5.4 57.7; Claude Opus 4.6 57.3; Gemini 54.2
NL2Repo42.735.9Claude 49.8; GPT-5.4 41.3
Terminal-Bench 2.0 Terminus-263.556.2Claude 65.4; Gemini 68.5
CyberGym68.748.3Claude 66.6; GPT-5.4 66.3
BrowseComp (without context management)68.062.0—
BrowseComp (with context management)79.375.9Claude 84.0; Gemini 85.9
MCP-Atlas Public71.869.2Claude 73.8; GPT-5.4 67.2

Long-horizon cases: More than 600 optimization iterations and 6,000+ tool calls in VectorDBBench, from about 3,547 to 21,500 QPS, with Recall constrained to ≥95%; a final geometric mean of about 3.6× speedup on KernelBench Level 3, compared with about 4.2× for Claude Opus 4.6; and a Linux desktop web app that ran for 8 hours, continuously adding features and fixing interactions.

Conclusion

Compared with GLM-5, GLM-5.1's most credible gains appear in “working longer” and “still finding structural improvements in optimization loops with feedback.” SWE-Bench Pro 58.4, CyberGym 68.7, and VectorDBBench 21.5k QPS support trying it as an engineering Agent; the mathematics and general-reasoning data, however, show that it is not the overall frontier leader.

Limitations

  • These are results self-reported by Z.ai and, unless otherwise noted, cannot be treated as independent controlled tests; the harness, judge, context trimming, and safety refusals all affect scores.

  • The 600-round/8-hour cases are individual showcase tasks; the full prompts, tool traces, failure samples, and costs have not all been disclosed.

  • Terminus-2, Claude Code, and best self-reported harness scores on Terminal-Bench are not interchangeable for comparison.

  • Some HLE/GPT/Claude comparisons carry asterisks; the official footnotes explicitly state that their sources or evaluation conditions differ. Do not treat the table as a set of fully homogeneous experiments.

Reproduction steps

  1. Fix the model version, harness, tools, timeout, context, parameters, and resource limits, and save the configuration file.

  2. First run public SWE-Bench Pro/NL2Repo/Terminal-Bench subsets, saving patches, tests, tool calls, and per-round metrics.

  3. Use the same feedback loop for VectorDBBench/KernelBench, clearly specifying constraints, baselines, submission frequency, and stopping conditions.

  4. Repeat code tasks multiple times at minimum; record tokens, cost, failures/fallbacks, and structural strategy changes, rather than only the final score.

  5. Compare GLM-5, Opus/GPT, and other models with the same harness, reporting pure reasoning, coding, tool use, and long-horizon efficiency separately.

Original evidence and data

The official source publishes the benchmark table, parameters/resources for each task type, the three long-horizon scenarios—VectorDBBench, KernelBench, and Linux desktop—and figures including 58.4 SWE-Bench Pro, 63.5 Terminus-2, and 68.7 CyberGym.

Source excerpt or observation (compliant short quote only)

The official source says GLM-5.1 “stays productive over longer sessions,” while also acknowledging that long-horizon optimization still faces local optima, coherence challenges across thousands of tool calls, and self-evaluation problems on tasks without metrics; this limits the scope of the conclusion.

What this supports

  • Supports the official long-horizon positioning and published benchmark framing.

What this does not support

  • Does not make the official results a bare-model ranking or your repository success rate.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Z.ai · Z.ai official · Original publication date 2026-04-07 · Site edit date 2026-09-20

Open original source

GLM-5.1

Compare GLM-5.1 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GLM-5.1 Explained: Long-Horizon Agents, Access, and Cost

A sourced GLM-5.1 overview covering its 200K context, 8-hour execution claim, Z.AI pricing snapshot, deployment boundaries, and a cautious pilot path.

Related reviews

GLM-5.1: Serenities AI's Self-Reported Benchmarks and the Boundaries of Independent ValidationSerenities AI’s 2026-03-29 evaluation separates an early Claude Code self-reported 45.3 from a later SWE-Bench Pro 58.4 and warns they are not the same test; its setup must be read as reported.GLM-5.1: Artificial Analysis Independent Intelligence Index and Inference Throughput BenchmarkThe Artificial Analysis GLM-5.1 Reasoning page collected 2026-08-20 records Intelligence Index 41 and 82.7 tokens/s, while noting verbosity and relatively high cost; this is an aggregated platform index.GLM-5.1: OpenCode Three-Model Industrial Webpage Benchmark and Real-World Capability BoundariesIn a single Reddit OpenCode industrial-dashboard comparison, GLM-5.1 produced the strongest visual UI and speed near DeepSeek V4 Pro but needed a second fix pass; Kubernetes YAML and 100k+ context exposed format/stability limits.GLM-5.1: Reddit LocalLLM Real-World Coding and Context ExperienceCommunity experiences describe GLM-5.1 as a cost-effective candidate for C++/everyday coding and long-running projects, but there is still significant disagreement over large monorepos, complex debugging, latency, and context stability. The provider and harness must be recorded..GLM-5.1: Long-horizon Agent and Claude Code ConfigurationGLM-5.1 should be configured as a “long-horizon engineering Agent”: provide ample context and output budget, clarify the role, tech stack, and acceptance criteria first, then let it loop through execution, compilation, testing, and iteration; in Claude Code, you can switch the model name directly to `GLM-5.1`..GLM-5.1: SGLang Heterogeneous Deployment and Interleaved Thinking ConfigurationLocal deployment of GLM-5.1 depends on the exact `transformers==5.3.0` version and SGLang parser configuration; coding Agent workflows must enable `Interleaved + Preserved Thinking` mode to prevent multi-turn forgetting..GLM-5.1: Claude Code Tool Discovery and System Role Compatibility WorkaroundWhen using GLM-5.1 in Claude Code or a multi-Agent framework, you must explicitly inject `tool_reference` parsing rules into the system prompt to prevent tool deadlocks, and intercept the `system` role in `messages[]` to avoid HTTP 422 errors..GLM-5.1: OpenCode Multi-Model Orchestration and Anti-Overthinking PromptEmbedding GLM-5.1 in a multi-model pipeline as a “high-value code executor,” together with a system prompt that enforces action, can effectively resolve overthinking deadlocks in Agents and YAML indentation defects..