Z.AI’s 2026-04-07 material claims up to 8 hours of sustained execution, 58.4 on SWE-Bench Pro, and 3.6× geometric-mean speedup on KernelBench Level 3; results depend on OpenHands/Terminus/Claude Code harnesses.
Z.ai · Read evidenceGLM-5.1 · Reviews and evidence
Which GLM-5.1 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
Serenities AI’s 2026-03-29 evaluation separates an early Claude Code self-reported 45.3 from a later SWE-Bench Pro 58.4 and warns they are not the same test; its setup must be read as reported.
Serenities AI · Read evidenceThe Artificial Analysis GLM-5.1 Reasoning page collected 2026-08-20 records Intelligence Index 41 and 82.7 tokens/s, while noting verbosity and relatively high cost; this is an aggregated platform index.
Artificial Analysis · Read evidenceFull reviews and related reading
Selected evidence
GLM-5.1: Z.ai's Official Long-Horizon Engineering Benchmarks and Reproduction Conditions
Z.AI’s 2026-04-07 material claims up to 8 hours of sustained execution, 58.4 on SWE-Bench Pro, and 3.6× geometric-mean speedup on KernelBench Level 3; results depend on OpenHands/Terminus/Claude Code harnesses.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.1; official release 2026-04-07; tasks include SWE-Bench Pro, KernelBench L3, and an 8-hour engineering loop; harness, context management, and parameters control comparability.
GLM-5.1: Serenities AI's Self-Reported Benchmarks and the Boundaries of Independent Validation
Serenities AI’s 2026-03-29 evaluation separates an early Claude Code self-reported 45.3 from a later SWE-Bench Pro 58.4 and warns they are not the same test; its setup must be read as reported.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.1; Serenities AI page dated 2026-03-29 with 04-07/04-10 updates; sample, tools, and baselines follow the page, and 45.3/58.4 cannot be combined.
GLM-5.1: Artificial Analysis Independent Intelligence Index and Inference Throughput Benchmark
The Artificial Analysis GLM-5.1 Reasoning page collected 2026-08-20 records Intelligence Index 41 and 82.7 tokens/s, while noting verbosity and relatively high cost; this is an aggregated platform index.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.1 Reasoning; Artificial Analysis collected 2026-08-20; index aggregates evaluations, with full items, provider, repeats, and hardware incomplete.
GLM-5.1: OpenCode Three-Model Industrial Webpage Benchmark and Real-World Capability Boundaries
In a single Reddit OpenCode industrial-dashboard comparison, GLM-5.1 produced the strongest visual UI and speed near DeepSeek V4 Pro but needed a second fix pass; Kubernetes YAML and 100k+ context exposed format/stability limits.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.1; one-shot OpenCode generation, source dated 2026-05-15; tasks were an industrial webpage, Kubernetes YAML, and 100k+ context; no repeated blind test.
All sources
All sources
GLM-5.1: Z.ai's Official Long-Horizon Engineering Benchmarks and Reproduction Conditions
Z.AI’s 2026-04-07 material claims up to 8 hours of sustained execution, 58.4 on SWE-Bench Pro, and 3.6× geometric-mean speedup on KernelBench Level 3; results depend on OpenHands/Terminus/Claude Code harnesses.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.1; official release 2026-04-07; tasks include SWE-Bench Pro, KernelBench L3, and an 8-hour engineering loop; harness, context management, and parameters control comparability.
GLM-5.1: Serenities AI's Self-Reported Benchmarks and the Boundaries of Independent Validation
Serenities AI’s 2026-03-29 evaluation separates an early Claude Code self-reported 45.3 from a later SWE-Bench Pro 58.4 and warns they are not the same test; its setup must be read as reported.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.1; Serenities AI page dated 2026-03-29 with 04-07/04-10 updates; sample, tools, and baselines follow the page, and 45.3/58.4 cannot be combined.
GLM-5.1: Artificial Analysis Independent Intelligence Index and Inference Throughput Benchmark
The Artificial Analysis GLM-5.1 Reasoning page collected 2026-08-20 records Intelligence Index 41 and 82.7 tokens/s, while noting verbosity and relatively high cost; this is an aggregated platform index.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.1 Reasoning; Artificial Analysis collected 2026-08-20; index aggregates evaluations, with full items, provider, repeats, and hardware incomplete.
GLM-5.1: OpenCode Three-Model Industrial Webpage Benchmark and Real-World Capability Boundaries
In a single Reddit OpenCode industrial-dashboard comparison, GLM-5.1 produced the strongest visual UI and speed near DeepSeek V4 Pro but needed a second fix pass; Kubernetes YAML and 100k+ context exposed format/stability limits.
Unverified: the original source could not be rechecked.
- Conditions
- Version GLM-5.1; one-shot OpenCode generation, source dated 2026-05-15; tasks were an industrial webpage, Kubernetes YAML, and 100k+ context; no repeated blind test.
GLM-5.1: Reddit LocalLLM Real-World Coding and Context Experience
Community experiences describe GLM-5.1 as a cost-effective candidate for C++/everyday coding and long-running projects, but there is still significant disagreement over large monorepos, complex debugging, latency, and context stability. The provider and harness must be recorded..
Unverified: the original source could not be rechecked.
- Source/version
- GLM-5.1: Reddit LocalLLM Real-World Coding and Context Experience; live status follows the source and was not reopened
- Task/sample
- personal-experience; tasks, samples, and repeats follow the disclosed portion
- Environment/harness
- Provider, client, parameters, and tool harness are not standardized
GLM-5.1
Compare GLM-5.1 in Tabbit
Model access, features, and permissions depend on your current client account.