Ox Alpha

Ox Alpha · Reviews and evidence

Which Ox Alpha conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

OpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.

OpenRouter · Read evidence

Cline says Ox Alpha and Fable both correctly fixed one real bug in its repository; Ox Alpha used about three times fewer output tokens and repeated less reasoning. This is an efficiency signal worth retesting, not a general capability ranking.

X · Read evidence

Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.

Reddit + GitHub experimental repository · Read evidence

Full reviews and related reading

Read the full analysis

Overview · English

Ox Alpha Explained: From Stealth Preview to GLM-5.3-Flash

Ox Alpha was the anonymous name for Z.ai GLM-5.3-Flash. Here are the verified specs, access boundaries, preview timeline and safe testing decision.

Selected evidence

OfficialVendor report

OpenRouter Record: Ox Alpha Specs, Availability, and Anonymous Provider

OpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.

SourceOpenRouter
Published2026-08-21
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha
Source
https://openrouter.ai/stealth/ox-alpha
Collection/review
2026-09-20; the dynamic source was not reopened
CodingAgentVisual generationReasoning
CommunityEditorial analysis

Cline Test: Ox Alpha and Fable Both Fixed a Real Repository Bug, with About 3x Less Output

Cline says Ox Alpha and Fable both correctly fixed one real bug in its repository; Ox Alpha used about three times fewer output tokens and repeated less reasoning. This is an efficiency signal worth retesting, not a general capability ranking.

SourceX
Published2026-08-25
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha
Source
https://x.com/cline/status/2091995642201842015
Collection/review
2026-09-20; the dynamic source was not reopened
CodingReasoningCost
CommunityIndependent measurement

Independent LiveCodeBench v6: Ox Alpha's Raw Pass@1

Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.

SourceReddit + GitHub experimental repository
Published2026-08-22
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha
Source
https://www.reddit.com/r/LLMDevs/comments/1vv4hmb/ox_alpha_livecodebench_v6/
Collection/review
2026-09-20; the dynamic source was not reopened
apiCodingAgent
CommunityIndependent measurement

Same Image, Same Prompt, Three Harnesses: Harness Effects on Ox Alpha

LeMi compared the DeepSeek, OMP, and OpenCode frameworks using the same Ox Alpha, the same *Interstellar* Ranger RF-31D reference image, and the same prompt, suggesting that a harness may substantially affect results, but the post does not publish a quantitative score.

SourceX
Published2026-08-26
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha
Source
https://x.com/LeMiMind/status/2092307287775846791
Collection/review
2026-09-20; the dynamic source was not reopened
CodingAgentVisual generationwriting

All sources

All sources

14 / 14
OfficialVendor report

OpenRouter Record: Ox Alpha Specs, Availability, and Anonymous Provider

OpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.

SourceOpenRouter
Published2026-08-21
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha
Source
https://openrouter.ai/stealth/ox-alpha
Collection/review
2026-09-20; the dynamic source was not reopened
CodingAgentVisual generationReasoning
CommunityEditorial analysis

Cline Test: Ox Alpha and Fable Both Fixed a Real Repository Bug, with About 3x Less Output

Cline says Ox Alpha and Fable both correctly fixed one real bug in its repository; Ox Alpha used about three times fewer output tokens and repeated less reasoning. This is an efficiency signal worth retesting, not a general capability ranking.

SourceX
Published2026-08-25
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha
Source
https://x.com/cline/status/2091995642201842015
Collection/review
2026-09-20; the dynamic source was not reopened
CodingReasoningCost
CommunityIndependent measurement

Independent LiveCodeBench v6: Ox Alpha's Raw Pass@1

Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.

SourceReddit + GitHub experimental repository
Published2026-08-22
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha
Source
https://www.reddit.com/r/LLMDevs/comments/1vv4hmb/ox_alpha_livecodebench_v6/
Collection/review
2026-09-20; the dynamic source was not reopened
apiCodingAgent
CommunityIndependent measurement

Same Image, Same Prompt, Three Harnesses: Harness Effects on Ox Alpha

LeMi compared the DeepSeek, OMP, and OpenCode frameworks using the same Ox Alpha, the same *Interstellar* Ranger RF-31D reference image, and the same prompt, suggesting that a harness may substantially affect results, but the post does not publish a quantitative score.

SourceX
Published2026-08-26
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha
Source
https://x.com/LeMiMind/status/2092307287775846791
Collection/review
2026-09-20; the dynamic source was not reopened
CodingAgentVisual generationwriting
CommunityPersonal experience

OpenCode Official Observation: 26T Ox Alpha Tokens in Four Days

OpenCode reports that Ox Alpha processed 26T tokens in four days, showing heavy real-world use of the preview but saying nothing by itself about model quality, individual quotas, or availability.

SourceX
Published2026-08-25
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-25
Method/client
Source-specific public post; client and provider conditions follow the source
CodingAgentCost
CommunityPersonal experience

OpenCode Go Entry: Free Period and Load Feedback

OpenCode announced Ox Alpha on OpenCode Go for six days of near-unlimited free use outside Go usage; public replies also report mid-run stops, roughly 20 tokens/s, and overload, so convenience and service stability must be evaluated separately.

SourceX
Published2026-08-21
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-21
Method/client
Source-specific public post; client and provider conditions follow the source
CodingAgentCostStability
CommunityPersonal experience

Aniruddha's Experience: Agent Tool Calls and Search Tasks

Aniruddha says Ox Alpha supports many agent tool calls and search-based features and that everything tested so far worked well; without tasks, traces, or a success definition, this is a tool-use experience awaiting reproduction.

SourceX
Published2026-08-26
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-26
Method/client
Source-specific public post; client and provider conditions follow the source
apiAgent
CommunityPersonal experience

Binx's Test: A Single-Sentence Fix in a Gauntlet

Binx says Ox Alpha fixed a gauntlet issue that DeepSeek, Qwen, MiniMax, and Sol had not fixed, using one sentence and about 10 seconds; this is a strong but non-reproducible single-case signal.

SourceX
Published2026-08-26
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-26
Method/client
Source-specific public post; client and provider conditions follow the source
Capability
CommunityPersonal experience

Matse's Test: Bug and Security Review of a One-Year Codebase

Matse says Ox Alpha found many bugs and security holes in a year's worth of code and fixed multiple problems in about three hours, but provides no sample or repair evidence; it is a candidate audit workflow, not a performance conclusion.

SourceX
Published2026-08-26
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-26
Method/client
Source-specific public post; client and provider conditions follow the source
Coding
CommunityEditorial analysis

Jonathan Turner: Ox Alpha Fingerprint Comparisons and Identity Boundaries

The article compares Ox Alpha with public GLM, Gemini, DeepSeek, Kimi, and MiMo using tokenizer behavior, video-token budgets, and server errors, strongly pointing to a GLM-family model on a Z.ai serving stack while explicitly stopping short of naming a product or developer.

SourceX Article
Published2026-08-25
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-25
Method/client
Source-specific public post; client and provider conditions follow the source
ReasoningCostStability
CommunityIndependent measurement

12-Prompt Stylometry Fingerprint Study: Ox Alpha's Similarity to GLM 5.3

Across 11 matched prompts, 7 reference models, and a deterministic 460-feature stylometry protocol, Ox Alpha was closest to GLM 5.3 on every prompt, but this indicates stylistic similarity under the test conditions rather than the identity of the weights or developer.

SourceReddit + GitHub experimental repository
Published2026-08-26
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-26
Method/client
Source-specific public post; client and provider conditions follow the source
Reasoningwriting
CommunityIndependent measurement

OpenCode Community Concurrency Experiment: About 40 Ox Alpha Agents

Ethan reports that running about 40 Ox Alpha agents simultaneously during the early free period produced about 7 tasks per worker per hour, 26.8 seconds P50 latency, and 100% traceable code citations; this is a single-user load observation, not a service SLA.

SourceReddit
Published2026-08-22
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-22
Method/client
Source-specific public post; client and provider conditions follow the source
CodingAgentCost
CommunityEditorial analysis

Ox Alpha vs. DeepSeek V4 Flash: Code Cleanup and Token-Use Experience Comparison

An OpenCode user says Ox Alpha cleaned up the results produced by DeepSeek V4 Flash in their project using about one-fifth as many tokens, while the same discussion includes counterexamples saying Ox was worse at logic, unsafe Rust, and assembly; the conclusion depends heavily on task type.

SourceReddit
Published2026-08-22
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-22
Method/client
Source-specific public post; client and provider conditions follow the source
CodingAgentCost
CommunityEditorial analysis

Same Prompt, Cross-Date Output Variance: Ox Alpha Version and Serving-Stack Uncertainty

Adit_Yah says the same prompt produced completely different code results three days apart, with duration increasing from 80 minutes to 330 minutes and more than 4,500 lines of code; this is better treated as a signal for reproducing routing, version, or sampling variance than as evidence of continual learning.

SourceX
Published2026-08-25
CollectedUnknown

Unverified: the original source could not be rechecked.

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-25
Method/client
Source-specific public post; client and provider conditions follow the source
Codingwriting

Ox Alpha

Compare Ox Alpha in Tabbit

Model access, features, and permissions depend on your current client account.