OpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.
OpenRouter · Read evidenceOx Alpha · Reviews and evidence
Which Ox Alpha conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
Cline says Ox Alpha and Fable both correctly fixed one real bug in its repository; Ox Alpha used about three times fewer output tokens and repeated less reasoning. This is an efficiency signal worth retesting, not a general capability ranking.
X · Read evidenceUnder a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.
Reddit + GitHub experimental repository · Read evidenceFull reviews and related reading
Selected evidence
OpenRouter Record: Ox Alpha Specs, Availability, and Anonymous Provider
OpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha
- Source
- https://openrouter.ai/stealth/ox-alpha
- Collection/review
- 2026-09-20; the dynamic source was not reopened
Cline Test: Ox Alpha and Fable Both Fixed a Real Repository Bug, with About 3x Less Output
Cline says Ox Alpha and Fable both correctly fixed one real bug in its repository; Ox Alpha used about three times fewer output tokens and repeated less reasoning. This is an efficiency signal worth retesting, not a general capability ranking.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha
- Source
- https://x.com/cline/status/2091995642201842015
- Collection/review
- 2026-09-20; the dynamic source was not reopened
Independent LiveCodeBench v6: Ox Alpha's Raw Pass@1
Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha
- Source
- https://www.reddit.com/r/LLMDevs/comments/1vv4hmb/ox_alpha_livecodebench_v6/
- Collection/review
- 2026-09-20; the dynamic source was not reopened
Same Image, Same Prompt, Three Harnesses: Harness Effects on Ox Alpha
LeMi compared the DeepSeek, OMP, and OpenCode frameworks using the same Ox Alpha, the same *Interstellar* Ranger RF-31D reference image, and the same prompt, suggesting that a harness may substantially affect results, but the post does not publish a quantitative score.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha
- Source
- https://x.com/LeMiMind/status/2092307287775846791
- Collection/review
- 2026-09-20; the dynamic source was not reopened
All sources
All sources
OpenRouter Record: Ox Alpha Specs, Availability, and Anonymous Provider
OpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha
- Source
- https://openrouter.ai/stealth/ox-alpha
- Collection/review
- 2026-09-20; the dynamic source was not reopened
Cline Test: Ox Alpha and Fable Both Fixed a Real Repository Bug, with About 3x Less Output
Cline says Ox Alpha and Fable both correctly fixed one real bug in its repository; Ox Alpha used about three times fewer output tokens and repeated less reasoning. This is an efficiency signal worth retesting, not a general capability ranking.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha
- Source
- https://x.com/cline/status/2091995642201842015
- Collection/review
- 2026-09-20; the dynamic source was not reopened
Independent LiveCodeBench v6: Ox Alpha's Raw Pass@1
Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha
- Source
- https://www.reddit.com/r/LLMDevs/comments/1vv4hmb/ox_alpha_livecodebench_v6/
- Collection/review
- 2026-09-20; the dynamic source was not reopened
Same Image, Same Prompt, Three Harnesses: Harness Effects on Ox Alpha
LeMi compared the DeepSeek, OMP, and OpenCode frameworks using the same Ox Alpha, the same *Interstellar* Ranger RF-31D reference image, and the same prompt, suggesting that a harness may substantially affect results, but the post does not publish a quantitative score.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha
- Source
- https://x.com/LeMiMind/status/2092307287775846791
- Collection/review
- 2026-09-20; the dynamic source was not reopened
OpenCode Official Observation: 26T Ox Alpha Tokens in Four Days
OpenCode reports that Ox Alpha processed 26T tokens in four days, showing heavy real-world use of the preview but saying nothing by itself about model quality, individual quotas, or availability.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-25
- Method/client
- Source-specific public post; client and provider conditions follow the source
OpenCode Go Entry: Free Period and Load Feedback
OpenCode announced Ox Alpha on OpenCode Go for six days of near-unlimited free use outside Go usage; public replies also report mid-run stops, roughly 20 tokens/s, and overload, so convenience and service stability must be evaluated separately.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-21
- Method/client
- Source-specific public post; client and provider conditions follow the source
Aniruddha's Experience: Agent Tool Calls and Search Tasks
Aniruddha says Ox Alpha supports many agent tool calls and search-based features and that everything tested so far worked well; without tasks, traces, or a success definition, this is a tool-use experience awaiting reproduction.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-26
- Method/client
- Source-specific public post; client and provider conditions follow the source
Binx's Test: A Single-Sentence Fix in a Gauntlet
Binx says Ox Alpha fixed a gauntlet issue that DeepSeek, Qwen, MiniMax, and Sol had not fixed, using one sentence and about 10 seconds; this is a strong but non-reproducible single-case signal.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-26
- Method/client
- Source-specific public post; client and provider conditions follow the source
Matse's Test: Bug and Security Review of a One-Year Codebase
Matse says Ox Alpha found many bugs and security holes in a year's worth of code and fixed multiple problems in about three hours, but provides no sample or repair evidence; it is a candidate audit workflow, not a performance conclusion.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-26
- Method/client
- Source-specific public post; client and provider conditions follow the source
Jonathan Turner: Ox Alpha Fingerprint Comparisons and Identity Boundaries
The article compares Ox Alpha with public GLM, Gemini, DeepSeek, Kimi, and MiMo using tokenizer behavior, video-token budgets, and server errors, strongly pointing to a GLM-family model on a Z.ai serving stack while explicitly stopping short of naming a product or developer.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-25
- Method/client
- Source-specific public post; client and provider conditions follow the source
12-Prompt Stylometry Fingerprint Study: Ox Alpha's Similarity to GLM 5.3
Across 11 matched prompts, 7 reference models, and a deterministic 460-feature stylometry protocol, Ox Alpha was closest to GLM 5.3 on every prompt, but this indicates stylistic similarity under the test conditions rather than the identity of the weights or developer.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-26
- Method/client
- Source-specific public post; client and provider conditions follow the source
OpenCode Community Concurrency Experiment: About 40 Ox Alpha Agents
Ethan reports that running about 40 Ox Alpha agents simultaneously during the early free period produced about 7 tasks per worker per hour, 26.8 seconds P50 latency, and 100% traceable code citations; this is a single-user load observation, not a service SLA.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-22
- Method/client
- Source-specific public post; client and provider conditions follow the source
Ox Alpha vs. DeepSeek V4 Flash: Code Cleanup and Token-Use Experience Comparison
An OpenCode user says Ox Alpha cleaned up the results produced by DeepSeek V4 Flash in their project using about one-fifth as many tokens, while the same discussion includes counterexamples saying Ox was worse at logic, unsafe Rust, and assembly; the conclusion depends heavily on task type.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-22
- Method/client
- Source-specific public post; client and provider conditions follow the source
Same Prompt, Cross-Date Output Variance: Ox Alpha Version and Serving-Stack Uncertainty
Adit_Yah says the same prompt produced completely different code results three days apart, with duration increasing from 80 minutes to 330 minutes and more than 4,500 lines of code; this is better treated as a signal for reproducing routing, version, or sampling variance than as evidence of continual learning.
Unverified: the original source could not be rechecked.
- Model/version
- Ox Alpha; exact snapshot follows the source
- Source date
- 2026-08-25
- Method/client
- Source-specific public post; client and provider conditions follow the source
Ox Alpha
Compare Ox Alpha in Tabbit
Model access, features, and permissions depend on your current client account.