GLM-5.3-Flash · Community source · Personal experience
Wenqi/Kevin shared an early hands-on impression and a DeepSWE subset result: Ox Alpha reached about **63%** at roughly 47K average output tokens per task. The author called it Pareto-frontier among open models and only slightly behind Grok 4.6 among closed models. This is a personal test and subjective comparison, not a standardized leaderboard. The post quotes OpenCode’s public Ox Alpha information (1M context, multimodal input, and zero data retention), but does not provide a full prompt set, task count, run configuration, or code. It also does not prove Ox Alpha’s GLM-5.3 Flash identity.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
Wenqi/Kevin shared an early hands-on impression and a DeepSWE subset result: Ox Alpha reached about 63% at roughly 47K average output tokens per task. The author called it Pareto-frontier among open models and only slightly behind Grok 4.6 among closed models. This is a personal test and subjective comparison, not a standardized leaderboard.
The post quotes OpenCode’s public Ox Alpha information (1M context, multimodal input, and zero data retention), but does not provide a full prompt set, task count, run configuration, or code. It also does not prove Ox Alpha’s GLM-5.3 Flash identity.
“Early community testing suggests that Ox Alpha is competitive on long-horizon coding tasks, but the 63% figure comes from a DeepSWE subset. It is not a full benchmark and cannot establish the model’s identity on its own.”
The result provides a traceable personal-test number. Its limitations are the subset and opaque test conditions; it should not be compared directly with Z.ai’s official GLM-5.3 DeepSWE v1.1 score.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X (Twitter) · Wenqi/Kevin (@winkeyh) · Original publication date 2026-08-21 · Site edit date 2026-09-20
Open original sourceGLM-5.3-Flash
Download the Tabbit client to check model access