Cline says Ox Alpha and Fable both correctly fixed one real bug in its repository; Ox Alpha used about three times fewer output tokens and repeated less reasoning. This is an efficiency signal worth retesting, not a general capability ranking.
Task: One real bug in the Cline repository.
Comparison models: Ox Alpha and Fable (the post does not provide versions, prices, or complete model IDs).
Outcome: The post says both models fixed the bug correctly.
Observation: Fable repeated “I found the root cause” seven times before editing; Ox Alpha stated it once and then wrote the fix.
The post does not disclose the bug number, repository commit, prompt, context, tool permissions, temperature, effort, timeout, repetition count, or evaluation script. It refers to thinking tokens and total output tokens but provides no raw count table, so “about three times less” must be recorded as Cline’s self-reported figure.
| Observation | Result in Cline’s post | Evidence boundary |
|---|---|---|
| Fix correctness | Ox Alpha and Fable both correctly fixed the same bug | One task; cannot represent an overall pass rate |
| Repeated reasoning | Fable repeated the root-cause statement 7 times; Ox Alpha once | Qualitative observation; full traces are not provided |
| Output efficiency | Ox Alpha used about 3x fewer output tokens for similar work | Self-reported total; no raw token counts or cost ledger |
| Early positioning | Cline also says early benchmarks put Ox Alpha slightly ahead of Fable and GPT | No benchmark, task set, or score is disclosed; do not cite as a ranking |
If the question is how much output an agent needs to complete the same coding fix, this post is a useful starting point for a reproducible test: use the same repository and bug, and record tool calls, thinking/output tokens, post-fix tests, and reviewer rework. The current evidence supports only “Cline reported one correct comparison with less repeated output,” not that Ox Alpha generally beats Fable or GPT.
For Tabbit users, Ox Alpha is a fit for a controlled, low-cost trial: start with a redacted repository, fixed tasks, and a reversible branch before widening the task set.
One real bug is too small a sample and has no independent verification.
Versions, context, effort, tool harness, and temperature for both models are undisclosed.
The “about 3x” figure has no raw token accounting and excludes latency, retries, and human review from total cost.
The “slightly ahead in early benchmarks” statement has no public scores or methodology and cannot be presented as leaderboard fact.
Ox Alpha’s anonymous provider, temporary free status, and data terms require separate verification.
Fix the Cline repository commit, the same bug, model versions, and permissions; clear prior sessions.
Run Ox Alpha and Fable separately, saving prompts, tool calls, thinking/output tokens, elapsed time, and the complete diff.
Let the test suite confirm the fix, then have an independent reviewer assess side effects and rework.
Expand to at least 20 varied real tasks and report means, medians, failure types, and confidence intervals; do not extrapolate the single-bug “about 3x” figure.
Ox Alpha