The article compares Ox Alpha with public GLM, Gemini, DeepSeek, Kimi, and MiMo using tokenizer behavior, video-token budgets, and server errors, strongly pointing to a GLM-family model on a Z.ai serving stack while explicitly stopping short of naming a product or developer.
Object: OpenRouter stealth/ox-alpha.
Input specification: 1,048,576-token context, text/image/video input, and text output; the public model card lists an anonymous Stealth provider.
Data sources: Measurements attributed to YFarmX, unclecode, Ben Davis, and Security Kid; the author explicitly says they were not rerun in this article.
Comparisons: Public GLM-5.3, GLM-5V-Turbo, Gemini, DeepSeek, Kimi, and MiMo.
| Measurement | Ox Alpha comparison | Interpretation boundary |
|---|---|---|
| Prompt tokens on 50 strings | 50/50 matches GLM-5.3; Gemini 11/50, DeepSeek 9/50, Kimi 9/50, MiMo 18/50 | Inherited community measurement; tokenizer match is not weight or product match |
| Video-token budget | Four clips match GLM-5V-Turbo at 296, 296, 884, and 1,064 respectively | n=4; encoder/budget similarity does not prove the same model |
| Server errors | The article says Java/Spring errors and business codes 1210/1214 are byte-identical to public GLM-5.3 errors | Shared serving stack is a hosting clue, not proof of weight ownership |
| Product-card comparison | Ox is 1M multimodal; public GLM-5.3 is 1M text-only, while GLM-5V-Turbo is 200K vision | “Same family” and “same public SKU” must remain separate |
The article's value is the reproducible numbers and explicit boundaries, not the claim that Ox Alpha is a named model. The stronger current conclusion is a GLM-style tokenizer/video encoder and a Z.ai PaaS hosting clue; the exact checkpoint, operator, and developer remain undisclosed.
Fix the 50 strings and tokenization method, query Ox Alpha and public models, and save prompt_tokens.
Fix the four video clips, dimensions, and frame rates; record visual-token or server-reported budgets and expand to far more than four samples.
Save raw HTTP errors, response headers, timestamps, and model versions; compare error bytes rather than screenshot similarity.
Score tokenizer, video budget, server stack, and product card separately. Update the identity conclusion only after a first-party claim, weight match, or larger-sample evidence appears.
The article says the measurements came from other people's work and were not rerun by its author.
Public model cards, anonymous-provider data, and community fingerprints involve inference; “GLM family” must not be rewritten as “GLM-5.3 Flash confirmed.”
The article's state is as of 2026-08-25; the server, model card, and anonymous preview can change.
Ox Alpha