GLM-5.3 · Media / benchmark · Platform telemetry
Exact-source rows trace provider numbers; without retesting they should not become an overall rank.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
BenchLM lists 17 GLM-5.3 scores whose original sources can be displayed, but does not assign an overall score or rank because there is no independent retest. For now, it is better understood as a traceable ledger of provider-reported data than as an independent leaderboard.
Suitable tasks: Checking GLM-5.3's published coding, Agent, and security benchmarks; quickly reviewing the source status of each number; and preparing a retest checklist for when the weights become available.
Unsuitable tasks: Treating the page's scores as an independent evaluation, an overall capability ranking, or the success rate of a real project.
Applicable model version: GLM-5.3 (the page labels it as released on 2026-08-14).
Applicable client, Agent, or API: Client-independent; the data comes from Z.AI's release table, and the model is currently available through Coding Plan/ZCode.
Recommended reasoning tier and parameters: BenchLM did not make any calls and did not publish parameters; the actual differences among low/high/max cannot be inferred from this data.
Data source: The BenchLM model page; its methodology page says public tables default to showing only exact-source rows, excluding generated or inferred scores.
Data refresh time: 2026-08-17.
Record status: GLM-5.3 has 17 displayable benchmark records; the Agentic and Coding categories each contain 8 records, plus 1 External signals record. Reasoning, Knowledge, Math, Multilingual, Multimodal, and Instruction Following are all marked as not measured.
Independence: The page explicitly marks these rows as Provider exact, and every original link points back to a Z.AI release article; there are no independently run inputs, configurations, seeds, or per-run results.
| Benchmark | GLM-5.3 | Page-recorded best verified comparison | Evidence status |
|---|---|---|---|
| Terminal-Bench 2.1 | 88.2% | GLM-5.3's current best verified row | Provider exact |
| Terminal-Bench 3.0 | 28.3% | GLM-5.3's current best verified row | Provider exact |
| DeepSWE v1.1 | 66.9% | GPT-5.6 Sol 72.7%, trails by 5.8 percentage points | Provider exact |
| NL2Repo | 58.0% | DeepSeek V4 Pro 0813 61.5%, trails by 3.5 percentage points | Provider exact |
| ProgramBench | 19.0% | Claude Opus 5 93.0%, a 74-percentage-point gap on the page | Provider exact |
| FrontierSWE | 78.1% | Kimi K3 81.2%, trails by 3.1 percentage points | Provider exact |
| SWE-Marathon v1.1 | 42.5% | GLM-5.3's current best verified row | Provider exact |
| PostTrainBench | 39.8% | GLM-5.3's current best verified row | Provider exact |
| CyberGym | 84.5% | Fugu Cyber 86.9%, trails by 2.4 percentage points | Provider exact |
| ExploitGym | Page-normalized value 15.0% | GPT-5.6 Sol 33.7%, trails by 18.7 percentage points | Provider exact |
| Toolathlon-Verified | 73.0% | Claude Opus 5 80.6%, trails by 7.6 percentage points | Provider exact |
| AutomationBench | 48.2% | GLM-5.3's current best verified row | Provider exact |
| Agents' Last Exam | 28.5% | Qwen3.8 Max 52.4%, trails by 23.9 percentage points | Provider exact |
| HLE with tools | 62.5% | Claude Opus 5 64.7%, trails by 2.2 percentage points | Provider exact |
| ExploitBench | 54.0% | Claude Mythos 5 78.0%, trails by 23.6 percentage points | Provider exact |
The BenchLM page also lists GDPval-AA v2 1769, but treats it as a provider-source record rather than an independently verified score; the page therefore remains at “Score pending / Not ranked.”
Facts that can be confirmed: GLM-5.3's published numbers are concentrated mainly in terminal coding, long-running Agent, and vulnerability discovery/exploitation tests; BenchLM can link each number to the corresponding Z.AI original.
Facts that cannot be confirmed: Whether these scores can be reproduced with other harnesses, inference stacks, budgets, model providers, or local weights; GLM-5.3's overall capability ranking; and whether the official token-efficiency claim translates into lower real-world task costs.
Implications for selection: In BenchLM's data, GLM-5.3 looks more like a model with "substantial coding/Agent evidence but obvious gaps in general-capability evidence." The page has not measured multimodality, math, knowledge, or instruction following, so coding results should not be extrapolated into all-purpose performance.
Boundary of the provider table: "Best verified row" only means the highest currently traceable source record in the directory; it does not mean BenchLM reran the benchmark and reproduced that score.
All GLM-5.3 rows have an evidence status of Provider exact, and their original links point to Z.AI; this is not an independent controlled evaluation.
The page does not publish GLM-5.3's request inputs, tool schema, time budget, sampling parameters, rollout count, hardware, or per-run results.
ExploitGym is normalized to 15.0% in the directory, while the Z.AI release table presents the result as the number of tasks completed in 2 hours/6 hours (105/130); the two representations cannot be treated as directly equivalent. Check BenchLM's normalization definition before comparing them.
Page data will change as the model directory is updated; this article records only the 2026-08-17 snapshot.
Open https://benchlm.ai/models/glm-5-3 and record the page's data date, the 17 source-displayable rows, and the “Score pending / Not ranked” status.
Follow the Z.AI original link for each record and verify the benchmark name, GLM-5.3 number, and comparison table; do not treat Google snippets or third-party retellings as results.
Open https://benchlm.ai/methodology and confirm that the directory includes only exact-source rows in its public table by default, then check the data refresh date.
For a genuinely independent retest, wait for downloadable weights or a measurable API. Then repeat the same tasks under a fixed repository and tool harness, recording inputs, context, effort, time budget, tool calls, each success/failure, and tokens per completed task.
The BenchLM model page explicitly states that GLM-5.3 has 17 benchmark rows whose sources can be displayed, but no public overall score or rank; all category scores on the page are pending.
The BenchLM methodology explicitly states that public tables default to exact-source rows only, and that generated benchmark values are not included in the evidence table; the dataset refresh date is 2026-08-17.
The individual original-source links for GLM-5.3 point back to the Z.AI release article https://z.ai/blog/glm-5.3; this article therefore treats BenchLM as a traceable directory, not a second independent scoring party.
“GLM-5.3 is tracked, but not publicly ranked yet.”
“Public BenchLM benchmark tables default to exact-source rows only.”
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
BenchLM (third-party model benchmark directory) · BenchLM editorial/data team (byline not disclosed) · Original publication date 2026-08-17 · Site edit date 2026-09-20
Open original sourceGLM-5.3
Download the Tabbit client to check model access