Wenqi/Kevin shared an early hands-on impression and a DeepSWE subset result: Ox Alpha reached about 63% at roughly 47K average output tokens per task. The author called it Pareto-frontier among open models and only slightly behind Grok 4.6 among closed models. This is a personal test and subjective comparison, not a standardized leaderboard.
The post quotes OpenCode’s public Ox Alpha information (1M context, multimodal input, and zero data retention), but does not provide a full prompt set, task count, run configuration, or code. It also does not prove Ox Alpha’s GLM-5.3 Flash identity.
“Early community testing suggests that Ox Alpha is competitive on long-horizon coding tasks, but the 63% figure comes from a DeepSWE subset. It is not a full benchmark and cannot establish the model’s identity on its own.”
The result provides a traceable personal-test number. Its limitations are the subset and opaque test conditions; it should not be compared directly with Z.ai’s official GLM-5.3 DeepSWE v1.1 score.
GLM-5.3-Flash