BenchLM lists 17 GLM-5.3 scores whose original sources can be displayed, but does not assign an overall score or rank because there is no independent retest. For now, it is better understood as a traceable ledger of provider-reported data than as an independent leaderboard.
Suitable tasks: Checking GLM-5.3's published coding, Agent, and security benchmarks; quickly reviewing the source status of each number; and preparing a retest checklist for when the weights become available.
Unsuitable tasks: Treating the page's scores as an independent evaluation, an overall capability ranking, or the success rate of a real project.
Applicable model version: GLM-5.3 (the page labels it as released on 2026-08-14).
Applicable client, Agent, or API: Client-independent; the data comes from Z.AI's release table, and the model is currently available through Coding Plan/ZCode.
Recommended reasoning tier and parameters: BenchLM did not make any calls and did not publish parameters; the actual differences among low/high/max cannot be inferred from this data.
Data source: The BenchLM model page; its methodology page says public tables default to showing only exact-source rows, excluding generated or inferred scores.
Data refresh time: 2026-08-17.
Record status: GLM-5.3 has 17 displayable benchmark records; the Agentic and Coding categories each contain 8 records, plus 1 External signals record. Reasoning, Knowledge, Math, Multilingual, Multimodal, and Instruction Following are all marked as not measured.
Independence: The page explicitly marks these rows as Provider exact, and every original link points back to a Z.AI release article; there are no independently run inputs, configurations, seeds, or per-run results.
| Benchmark | GLM-5.3 | Page-recorded best verified comparison | Evidence status |
|---|---|---|---|
| Terminal-Bench 2.1 | 88.2% | GLM-5.3's current best verified row | Provider exact |
| Terminal-Bench 3.0 | 28.3% | GLM-5.3's current best verified row | Provider exact |
| DeepSWE v1.1 | 66.9% | GPT-5.6 Sol 72.7%, trails by 5.8 percentage points | Provider exact |
| NL2Repo | 58.0% | DeepSeek V4 Pro 0813 61.5%, trails by 3.5 percentage points | Provider exact |
| ProgramBench | 19.0% | Claude Opus 5 93.0%, a 74-percentage-point gap on the page | Provider exact |
| FrontierSWE | 78.1% | Kimi K3 81.2%, trails by 3.1 percentage points | Provider exact |
| SWE-Marathon v1.1 | 42.5% | GLM-5.3's current best verified row | Provider exact |
| PostTrainBench | 39.8% | GLM-5.3's current best verified row | Provider exact |
| CyberGym | 84.5% | Fugu Cyber 86.9%, trails by 2.4 percentage points | Provider exact |
| ExploitGym | Page-normalized value 15.0% | GPT-5.6 Sol 33.7%, trails by 18.7 percentage points | Provider exact |
| Toolathlon-Verified | 73.0% | Claude Opus 5 80.6%, trails by 7.6 percentage points | Provider exact |
| AutomationBench | 48.2% | GLM-5.3's current best verified row | Provider exact |
| Agents' Last Exam | 28.5% | Qwen3.8 Max 52.4%, trails by 23.9 percentage points | Provider exact |
| HLE with tools | 62.5% | Claude Opus 5 64.7%, trails by 2.2 percentage points | Provider exact |
| ExploitBench | 54.0% | Claude Mythos 5 78.0%, trails by 23.6 percentage points | Provider exact |
The BenchLM page also lists GDPval-AA v2 1769, but treats it as a provider-source record rather than an independently verified score; the page therefore remains at “Score pending / Not ranked.”
Facts that can be confirmed: GLM-5.3's published numbers are concentrated mainly in terminal coding, long-running Agent, and vulnerability discovery/exploitation tests; BenchLM can link each number to the corresponding Z.AI original.
Facts that cannot be confirmed: Whether these scores can be reproduced with other harnesses, inference stacks, budgets, model providers, or local weights; GLM-5.3's overall capability ranking; and whether the official token-efficiency claim translates into lower real-world task costs.
Implications for selection: In BenchLM's data, GLM-5.3 looks more like a model with "substantial coding/Agent evidence but obvious gaps in general-capability evidence." The page has not measured multimodality, math, knowledge, or instruction following, so coding results should not be extrapolated into all-purpose performance.
Boundary of the provider table: "Best verified row" only means the highest currently traceable source record in the directory; it does not mean BenchLM reran the benchmark and reproduced that score.
All GLM-5.3 rows have an evidence status of Provider exact, and their original links point to Z.AI; this is not an independent controlled evaluation.
The page does not publish GLM-5.3's request inputs, tool schema, time budget, sampling parameters, rollout count, hardware, or per-run results.
ExploitGym is normalized to 15.0% in the directory, while the Z.AI release table presents the result as the number of tasks completed in 2 hours/6 hours (105/130); the two representations cannot be treated as directly equivalent. Check BenchLM's normalization definition before comparing them.
Page data will change as the model directory is updated; this article records only the 2026-08-17 snapshot.
Open https://benchlm.ai/models/glm-5-3 and record the page's data date, the 17 source-displayable rows, and the “Score pending / Not ranked” status.
Follow the Z.AI original link for each record and verify the benchmark name, GLM-5.3 number, and comparison table; do not treat Google snippets or third-party retellings as results.
Open https://benchlm.ai/methodology and confirm that the directory includes only exact-source rows in its public table by default, then check the data refresh date.
For a genuinely independent retest, wait for downloadable weights or a measurable API. Then repeat the same tasks under a fixed repository and tool harness, recording inputs, context, effort, time budget, tool calls, each success/failure, and tokens per completed task.
The BenchLM model page explicitly states that GLM-5.3 has 17 benchmark rows whose sources can be displayed, but no public overall score or rank; all category scores on the page are pending.
The BenchLM methodology explicitly states that public tables default to exact-source rows only, and that generated benchmark values are not included in the evidence table; the dataset refresh date is 2026-08-17.
The individual original-source links for GLM-5.3 point back to the Z.AI release article https://z.ai/blog/glm-5.3; this article therefore treats BenchLM as a traceable directory, not a second independent scoring party.
“GLM-5.3 is tracked, but not publicly ranked yet.”
“Public BenchLM benchmark tables default to exact-source rows only.”
GLM-5.3