OpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.
Entry point: OpenRouter's stealth/ox-alpha model record.
Provider: The page shows one provider named Stealth; OpenRouter says it only routes requests and is not the developer, owner, or provider.
Page state: Free preview; the pricing table shows $0/M tokens for both input and output.
Dynamic page metrics (at collection): The provider table showed 5.78 seconds P50 latency and 24 tokens/s throughput; the page's one-week aggregate separately showed 31 tokens/s average P50 throughput and 4.44 seconds average P50 latency. These have different scopes and should not be merged into one benchmark score.
Context limit: 1,048,576 tokens (about 1M).
Maximum output: 131,072 tokens.
Input modalities: Text, image, and video.
Positioning: A reasoning model for long-horizon software engineering, complex reasoning, and workflows combining text with visual context.
Not disclosed: Prompts, temperature, effort, tool harness, repeat count, task set, and failure traces are not provided on the page.
| Item | Verifiable page information | Interpretation boundary |
|---|---|---|
| Input price | $0/M | Current preview price, not a long-term commitment |
| Output price | $0/M | Current preview price, not a long-term commitment |
| Context | 1,048,576 tokens | Capacity limit, not proof of whole-repository recall |
| Maximum output | 131,072 tokens | Server limit, not an expectation for every response |
| P50 latency | 5.78 seconds (provider table) | Dynamic observation; record collection time |
| P50 throughput | 24 tokens/s (provider table) | Different scope from the one-week 31 tokens/s summary |
| 3-day availability | 98.78% | OpenRouter service observation, not a capability benchmark |
This is an anonymous preview route with relatively clear headline specifications but limited capability evidence. It is suitable for low-cost coding, long-context, and agent prototypes when data is non-sensitive; the 1M context, free price, or current latency alone cannot establish superiority over GPT, Claude, GLM, or other models.
The OpenRouter page says prompts and completions are retained by the provider but not used for training; it also defers other use to the Stealth Model Terms. Recheck the applicable terms before use and do not treat an anonymous preview as a privacy or availability SLA.
No developer name, weights, version snapshot, or public technical report.
No SWE-bench, Arena, or unified-harness score is published; the current metrics are routing/service observations.
Latency, throughput, and availability change with time and load; live observations are not fixed performance guarantees.
Free access is a preview condition; future pricing, limits, and availability are unknown.
The page's retention statement and the general Stealth terms should be interpreted using the stricter boundary.
Open stealth/ox-alpha in Tabbit and record the page date, provider, pricing, and service metrics.
Use a redacted repository and fix 10–20 coding and long-context tasks, saving the full prompt, tool permissions, model ID, and response trace.
Repeat at least three times and report first-pass success, post-fix tests, latency, output tokens, retries, and reviewer rework.
Compare with a known-snapshot baseline using the same tasks, harness, and permissions before drawing conclusions.
Ox Alpha