In Arena.ai's published Code Arena: WebDev real-world evaluation, Claude Sonnet 5.5 (High) ranks fourth with a score of 1699 and enters the cost-efficiency Pareto frontier at a blended price of $8/Mtoken; it improves by 159 points over Sonnet 5 (High), but the post does not publish the complete sample, scoring method, or runtime configuration.
Suitable tasks: Real code generation and web development tasks represented by WebDev; it can serve as an external signal when deciding whether Sonnet 5.5 High belongs in a candidate model pool.
Unsuitable tasks: Extrapolating a single Code Arena sub-leaderboard directly to cross-file refactoring, long-horizon Agents, production reliability, or other programming languages.
Applicable model version: Claude Sonnet 5.5 (High), the Arena evaluation snapshot in the X post dated 2026-09-30.
Applicable client, agent, or API: Arena.ai Code Arena: WebDev; the post does not publish the model provider, complete harness, per-task prompts, or API parameters.
Recommended reasoning effort and parameters: The post reports only the High level. For a repeat evaluation, fix High, the model snapshot, tool permissions, context, output limit, and billing basis.
Evaluation platform: Arena.ai's Code Arena: WebDev, which the post describes as real-world results and reports with both overall and subdomain rankings.
Model configuration: Claude Sonnet 5.5 (High); the comparisons are Sonnet 5 (High), Claude Fable 5.1 (Max), and GPT-6 Astra (Max).
Pricing basis: The post uses a blended price of $8 per Mtoken and says Sonnet 5.5 High is 80% cheaper than Fable 5.1 Max, ranked third overall, and GPT-6 Astra Max, ranked second.
Original post status: The post body was visible on X, and the English content was checked after switching to “Show original.” Its Pareto chart prints the Code Arena: WebDev leaderboard URL, but the post does not include a clickable leaderboard link.
Time snapshot: X shows a publication time of 2026-09-30 02:15 (as displayed in the page's local time). The post had about 92,000 views at collection time; engagement counts change dynamically.
| Metric | Claude Sonnet 5.5 (High) | Comparison or change |
|---|---|---|
| Code Arena: WebDev overall score | 1699 points | Ranked 4th |
| Sonnet 5 (High) | 1540 points | Ranked 37th; Sonnet 5.5 High improves by +159 points |
| Reference-Based Design | Ranked 4th | Sonnet 5 High ranked 38th |
| Simulations | Ranked 4th | Sonnet 5 High ranked 37th |
| Gaming | Ranked 4th | Sonnet 5 High ranked 36th |
| Blended price | $8/Mtoken | Arena.ai says its cost efficiency reshapes the WebDev Pareto frontier |
This first-hand Arena.ai update supports a bounded conclusion: Claude Sonnet 5.5 High enters the top four on the Code Arena: WebDev real-world task leaderboard and shows a clear ranking and score improvement over Sonnet 5 High; under the blended pricing basis given in the post, it also has good cost efficiency. The conclusion covers only the WebDev leaderboard and the High configuration, and cannot replace a test suite using your own codebase, cross-file changes, test repairs, and long-horizon Agent tasks.
The X post does not publish the complete Code Arena: WebDev task set, sample size, number of evaluators, Elo or other score formula, randomization method, or confidence intervals.
1699 points and fourth place are a dynamic leaderboard snapshot; new models, sample updates, or leaderboard recalculation can change the ranking.
The post uses a blended $8/Mtoken cost basis, but does not publish the input/output token ratio, cache hits, provider, tool calls, or per-task cost distribution. It therefore cannot be used directly to budget production costs.
The WebDev leaderboard does not represent terminal Agents, cross-file refactoring, code security, test coverage, non-web programming, or multi-turn project maintenance capabilities.
The post's statement that it is “80% cheaper” is a comparison against Fable 5.1 Max and GPT-6 Astra Max in the overall comparison table. It does not provide the complete price calculation and should be checked using the same provider and task set.
Fix the currently available Code Arena: WebDev data version, model snapshot, High level, and billing provider on Arena.ai, and record the leaderboard retrieval time.
Supplement the evaluation with your own WebDev task set covering reference implementations, simulations/interactions, games, and cross-file web tasks; save the complete prompts, repository versions, tool traces, and generated results.
Compare Sonnet 5.5 High, Sonnet 5 High, and other candidate models in the same harness, with timeout, context, output limit, temperature, and retry strategy held constant.
Calculate task success rate, human preference, test pass rate, repair rounds, input/output tokens, latency, and actual cost separately; do not treat 1699 points as your own rerun result.
Archive leaderboard rankings, scores, and prices with the collection time; when the leaderboard updates, preserve the old snapshot and explain the reason for the change.
The Arena.ai post summarizes Sonnet 5.5 High as the fourth-ranked WebDev model with “1699 pts” and emphasizes its cost efficiency; these numbers apply only to that leaderboard and the 2026-09-30 snapshot.
Claude Sonnet 5.5