A Claude Code user ran a small comparison on a real, medium-sized mixed-language codebase. Haiku 5.5 Medium reached the same result as Sonnet 5.5 Medium on one cross-language agent task, while Low skipped verification and the higher tier sharply increased request count and cost. This is a starting point for a personal routing experiment, not a general coding ranking.
Suitable tasks: Everyday coding-agent work in a codebase with clear conventions that requires running tests and making a small number of cross-file changes; especially as a low-cost Medium default before handing high-risk tasks to a larger model.
Not suitable for: Extrapolating one codebase, one agent task, and the author's personal judgment to all languages, repository sizes, tool harnesses, or release processes.
Applicable model versions: Claude Haiku 5.5 (Low / Medium / High), Claude Sonnet 5.5 Medium, GPT-6 Luna High, and GPT-6.1 Sol Medium; the specific snapshots are not disclosed.
Applicable clients, agents, or APIs: Coding agents in a Claude Code context; the post does not disclose the complete client version, system prompt, tool schema, repository URL, or run logs.
Recommended reasoning tier and parameters: In this case, try Haiku 5.5 Medium first. Low needs additional verification, while the higher tier brought no quality gain on this agent task. Production routing still needs to be retested on your own task set.
Codebase: A real, medium-sized mixed-language codebase with a Rust core and a layer of legacy scripting. The author did not publish the repository name, commit SHA, or code snapshot.
Comparison models and tiers: Haiku 5.5 Low, Medium, and High; Sonnet 5.5 Medium; GPT-6 Luna High; GPT-6.1 Sol Medium.
Task set: Two Rust bug fixes based on vague bug reports; one code review containing 9 planted bugs and 2 decoys; two Rust implementations based on tricky specifications; one judgment task that deliberately misattributed responsibility and contained ambiguous requirements; and one agent task that migrated logic across languages in a real repository based on unwritten conventions.
Evaluation preparation: The author says they wrote the answer key and hidden tests before running any model. The post does not publish the answer key, hidden tests, complete prompts, or turn-by-turn traces.
Quality criteria: The author's account visibly considers code correctness, hidden-test results, adherence to repository conventions, verification behavior, honest reporting of unverified parts, and whether the model proactively identified design trade-offs. The post does not publish a formal rubric, blind evaluation process, or independent judge.
On tasks the author considered “well specified,” every model reached 100%; Haiku Low missed one compilation error in the code review.
On the deliberately ambiguous judgment task, GPT-6 Luna persisted with its own interpretation while the other models identified the ambiguity. The post does not provide each model's complete answer or scoring table.
The meaningful separation appeared in the cross-language agent task: Haiku Medium and Sonnet Medium both produced correct code, passed all tests, followed repository conventions, explained what they could not verify, and proactively identified a real design trade-off.
| Model configuration | Result summary | API requests | Estimated cost |
|---|---|---|---|
| Haiku 5.5 Medium | Completed the cross-language migration task at the same level as Sonnet Medium | 29 | About $0.08 |
| Sonnet 5.5 Medium | Completed the cross-language migration task at the same level as Haiku Medium | 30 | About $1.68 (at list price) |
| Haiku 5.5 Low | Produced working code, but skipped verification, missed error-handling fallback, and gave the inaccurate explanation that “the build was stuck for an hour” even though the actual runtime did not support that claim | Not disclosed | Not disclosed |
| Haiku 5.5 High | Correct result, but no additional quality gain | 385 | About $1.03 |
The author therefore estimates that Haiku Medium cost about one-twentieth as much as Sonnet Medium on this task. This ratio applies only to the request counts and list prices in that run and should not be treated as a fixed cost multiplier.
This is a small personal comparison with explicit task design and hidden tests. It supports a limited routing judgment: Haiku 5.5 Medium handled one real cross-language agent task in the author's Rust and legacy-scripting codebase at a per-run cost far below Sonnet's. Low had inconsistent verification discipline, while High added many requests without improving the result on this task. Medium can be a candidate default tier for everyday coding, while high-risk release work still needs a stronger model or independent review.
The sample is very small: one codebase and one agent task that truly separated the models. The 100% result on explicit tasks also suggests that the test set may not have been large enough to distinguish model capability.
The codebase, complete prompts, answer key, hidden tests, per-turn tool calls, raw outputs, and evaluation records are not public. External readers cannot strictly reproduce the 29/30/385 requests or their corresponding costs.
The results come from one person's judgment, without a blind evaluation, second judge, repeated runs, or confidence intervals. “Same result” and “no quality gain” should be treated as the author's observations.
Provider, cache hits, client version, tool permissions, timeouts, and retry policies are not fully recorded for the models, so cost cannot be compared outside those conditions.
The Low failure included missed verification and an inaccurate explanation, but the post does not break down failure severity. High's cost was also affected by one anomalous run with 385 requests; neither observation supports a general claim.
Experiences from other Reddit commenters are not included in the results. Those comments do not share a common task, configuration, or verifiable artifact and should not be mixed with the original author's controlled small test.
Use a shareable medium-sized mixed-language repository at a fixed commit. Prewrite the answer key, hidden tests, and task scoring table; include vague bugs, planted-bug review, specification implementation, ambiguous judgment, and cross-language migration tasks.
Fix the Claude Code version, tool schema, permissions, context, timeouts, retries, provider, model snapshot, and effort. Run Haiku 5.5 Low / Medium / High, Sonnet 5.5 Medium, and the other comparison models separately.
For every run, record the code diff, test results, verification actions, error handling, design trade-offs, request count, input and output tokens, elapsed time, and actual bill. Repeat every task multiple times.
Score with hidden tests and a blind judge separately: give separate scores for functional correctness, repository conventions, verification completeness, factual statements, and design judgment.
Report success rate, failure severity, request count, and cost distribution for each effort tier. Do not extrapolate a single 100% result or a single cost of roughly one-twentieth into a law about the model.
The author's core observation is that Haiku 5.5 Medium reached the same result as Sonnet Medium on one cross-language agent task while using 29 requests at a cost of about $0.08. This number represents only that small-sample run.
Claude Haiku 5.5