Artificial Analysis's independent evaluation places Claude Sonnet 5.5 at number 2 in the Artificial Analysis Intelligence Index with a score of 56 at max effort, and finds it close to Opus 5.5 on terminal Agent and knowledge work tasks, at the cost of the highest output token usage measured by the site and a higher per-task cost.
Suitable tasks: Agent tasks that require terminal tools, code execution, real workflows, and longer reasoning chains; Terminal-Bench 4.0, AA-Briefcase, GDPval-AA, and AutomationBench-AA are the most relevant directions.
Unsuitable tasks: Production workflows that are highly sensitive to factual accuracy, scientific reasoning, or a cost ceiling without additional budget and result verification.
Applicable model version: Claude Sonnet 5.5 pre-release deployment (the snapshot used in the Artificial Analysis article); the article is not the final regression result for the public release.
Applicable client, agent, or API: The Artificial Analysis Intelligence Index evaluation environment; the article does not publish copyable API call scripts or a complete tool schema.
Recommended reasoning effort and parameters: The article compares five levels, low, medium, high, xhigh, and max, with Anthropic's default fallback enabled. For a repeat evaluation, fix effort, fallback, tool set, timeout, and token limit at each level.
Evaluation suite: The Artificial Analysis Intelligence Index, covering terminal agents, knowledge work, automation, scientific coding, factual knowledge, and long-context tasks.
Model input and context: The article lists a 1M-token context window and text and image input. These are model specifications, not the context length used in every evaluation run.
Comparison models: Claude Opus 5.5, Claude Sonnet 5, GPT-6 Astra, and GPT-6 Sol, among others; different tasks use their own evaluation harnesses, so all numbers should not be treated as one overall score on a single task set.
Reasoning levels: Sonnet 5.5 at low, medium, high, xhigh, and max, with Anthropic's default fallback enabled.
Fallback: Artificial Analysis observed fallback on about 0.1% of Intelligence Index tasks. The article says these cases all fell back to Sonnet 5 and occurred mainly on Terminal-Bench 4.0.
Pre-release limitation: The tested deployment contained a bug that reduced the quality of structured outputs requests. Anthropic fixed the issue before the public release and expected the final results to change very little or to show that the earlier results had been slightly underestimated; Artificial Analysis said it would rerun the relevant evaluations.
Pricing context: The article records Sonnet 5.5 and Sonnet 5 at the same price: $2 input, $10 output, $2.5 cache write, and $0.2 cache read per million tokens.
| Evaluation | Claude Sonnet 5.5 | Comparison or note |
|---|---|---|
| Artificial Analysis Intelligence Index (max) | 56 | Second only to Claude Opus 5.5 (max) at 58; 18 points higher than Sonnet 5 (max) |
| Terminal-Bench 4.0 (max) | 64% | Opus 5.5 scored 60%; GPT-6 Astra (xhigh) scored 60% |
| Terminal-Bench-Science (max) | 53% | Behind GPT-6 Astra and Opus 5.5; this evaluation was not included in the Intelligence Index at the time |
| AA-Briefcase | 1811 Elo | Opus 5.5 scored 1822 Elo, nearly tied |
| GDPval-AA | 1844 Elo | Opus 5.5 scored 1846 Elo, nearly tied |
| AutomationBench-AA | 71% | Opus 5.5 scored 70%; this was the headline score listed in the article |
Output tokens: At max, each Intelligence Index task used about 193k output tokens, the highest amount Artificial Analysis had measured at the time; this was about 60% higher than Opus 5.5 (max) or Sonnet 5 (max), and about 7 times GPT-6 Astra (max).
Per-task cost: The article puts Sonnet 5.5 at about $7.60 per task, roughly 50% higher than Sonnet 5's per-task cost. In the tradeoff between Intelligence and Cost per Task, high effort is the most competitive, though it remains slightly behind GPT-6 Sol at a similar cost. Lower effort levels need separate quality and cost comparisons.
AA-Omniscience factual accuracy: 54% for Sonnet 5.5 and 66% for Opus 5.5.
AA-Omniscience hallucination rate: 47% for Sonnet 5.5 and 59% for Opus 5.5; lower is better. This metric should not be treated directly as the rate of factually correct answers.
Humanity's Last Exam and SciCode: The article says Sonnet 5.5 was about 6 points below Opus 5.5, but does not provide a complete per-item score table for either evaluation in the body.
In Artificial Analysis's measurements, Sonnet 5.5 moves the Sonnet series into a range close to Opus 5.5 on terminal use and Agent workflows: it scores 64% on Terminal-Bench 4.0, while AA-Briefcase, GDPval-AA, and AutomationBench-AA are also close to or slightly above Opus 5.5. The main tradeoff is extremely high output token usage, so model selection should evaluate per-task cost, latency, and context budget separately from capability. Factual accuracy and scientific reasoning remain below Opus 5.5, so a high overall index ranking does not remove the need for fact checking.
The article measures a pre-release deployment, and the structured outputs bug was fixed before public release; the relevant results need to be checked against Artificial Analysis's rerun.
Artificial Analysis does not publish each task's complete prompt, sample size, temperature, tool schema, turn-by-turn trace, or raw output in the article. The conclusions and main metrics can therefore be checked, but the complete evaluation cannot be rerun from the article alone.
Intelligence Index, Terminal-Bench, GDPval-AA, AutomationBench-AA, and AA-Omniscience use different task sets and scoring methods; their scores cannot be compared directly as a single success rate.
The approximately $7.60 per-task cost is from the collection snapshot; model pricing, token usage, provider routing, and cache hits can change the actual cost.
Fallback rate, token usage, and leaderboard position all depend on the model deployment and Artificial Analysis harness at the time, so they cannot be unconditionally extrapolated to other API providers, clients, or self-built Agents.
Fix the public Claude Sonnet 5.5 API snapshot, provider, effort level, default fallback behavior, and token limit.
Using the same Agent harness, separately rerun Terminal-Bench 4.0, Terminal-Bench-Science, AA-Briefcase-style knowledge work, GDPval-AA-style office tasks, and AutomationBench-AA-style SaaS workflows.
For each task, save the input, tool calls, output tokens, reasoning tokens, latency, failure type, human corrections, and cost; repeat each task several times at minimum to estimate variance.
Separately measure AA-Omniscience-style factual accuracy and hallucination rate, and report factual errors separately from safe refusals.
Report the rerun results alongside the snapshot metrics in this article, including 56, 64%, 1811 Elo, 1844 Elo, and 71%; do not label your own results as the original Artificial Analysis scores.
The article's core observation is that Sonnet 5.5 approaches Opus 5.5 at the max level while using about 193k output tokens per task; this conclusion must be cited together with the corresponding evaluation harness and the pre-release version limitation.
Claude Sonnet 5.5