The most valuable part of this full evaluation is not the “94.6% of Opus” headline, but the distinction it draws between the early Claude Code self-reported result of 45.3 and the later SWE-Bench Pro update of 58.4, while clearly warning that the two must not be treated as the same test.
Suitable tasks: Assessing GLM-5.1's coding positioning, the significance of its price and open-source deployment, and the risk checklist for cases where independent validation is still needed.
Unsuitable tasks: Using the early 45.3 or later 58.4 as a single number to directly decide on a production replacement; the article itself notes that the early result lacks complete methodological details.
Applicable model versions: Materials from the initial GLM-5.1 release period and the 2026-04 updates.
Applicable clients, agents, or APIs: Claude Code harness, Z.ai API, Cline, and others; the article does not disclose a reproducible, unified invocation configuration.
Recommended reasoning level and parameters: Not publicly disclosed / cannot be verified; use Z.ai's current official benchmarks and documentation as the source of truth.
The early coding scores use Claude Code as the evaluation framework: GLM-5.1 45.3, GLM-5 35.4, and Claude Opus 4.6 47.9.
The article explicitly states that the early evaluation did not fully disclose its scoring method and used the Claude Code harness, so it cannot be directly compared across other benchmarks.
The 2026-04 update cites Z.ai's new SWE-Bench Pro 58.4, CyberGym 68.7, Terminal-Bench 63.5, NL2Repo 42.7, and other results; their parameters and resource conditions should be taken from Z.ai's official page.
| Model | Coding score | Relative to Opus 4.6 |
|---|---|---|
| Claude Opus 4.6 | 47.9 | 100% |
| GLM-5.1 | 45.3 | 94.6% |
| GLM-5 | 35.4 | 73.9% |
| Benchmark | GLM-5.1 | GPT-5.4 | Claude Opus 4.6 |
|---|---|---|---|
| SWE-Bench Pro | 58.4 | 57.7 | 57.3 |
| CyberGym | 68.7 | Not publicly disclosed | 66.6 |
| Terminal-Bench 2.0 | 63.5 (Claude Code scaffold separately reported at 66.5) | Not publicly disclosed | Not publicly disclosed |
| HLE | 31.0 | 39.8 | 36.7 |
| GPQA-Diamond | 86.2 | 92.0 | 91.3 |
| AIME 2026 | 95.3 | 98.7 | Not publicly disclosed |
| NL2Repo | 42.7 | Not publicly disclosed | Not publicly disclosed |
The evidence from Serenities AI supports three points: GLM-5.1's engineering and coding direction merits attention; the early 45.3 result was self-reported by Z.ai and strongly dependent on the harness; and the later figures such as 58.4 represent a different, updated evaluation. The article also points out that GLM has practical value in terms of cost, MIT licensing, and data sovereignty, but that its general reasoning and performance on complex, large codebases still require hands-on testing.
Much of the article is news-style secondary compilation and cannot substitute for the original benchmark.
The early 45.3 and later 58.4 come from different evaluation setups and time points, so they cannot be turned into a “model improvement percentage.”
The background information in the article about the IPO, hardware, pricing, and some independent validation has no direct causal relationship to model quality and should be verified separately.
The article does not disclose the complete prompt, tool schema, number of repetitions, or raw outputs; its category is comparison rather than reproducible-test.
Treat the early Claude Code score and the current SWE-Bench Pro result as two independent experiments, and do not mix their data.
Reproduce the parameters of the currently public benchmarks directly from Z.ai's official footnotes, then record the version and date.
For real projects, select small fixes, medium-sized refactors, debugging, and large-repository tasks, and run A/B tests with a fixed harness and acceptance criteria.
Also record token cost, latency, context loss, tool errors, and manual corrections to avoid substituting a single score for production evidence.
The article provides the complete early 45.3/47.9/35.4 table and gives 58.4 SWE-Bench Pro and other benchmarks in the update section, while explicitly stating that the early figures were “entirely self-reported by Z.ai” and awaiting third-party review.
The evaluation's central warning is “Treat the 94.6% figure as a promising preliminary claim, not an established fact”; this is a more appropriate framing of the material's boundaries than treating the headline number as a conclusion.
GLM-5.1