Claude Opus 5.5 is the easiest Opus decision in a while: switch your default Opus work to it. Anthropic released it on September 22, 2026 at $4 per million input and $20 output tokens — 20% below Opus 5 — and pitched it as performing "at the level of Claude Fable 5.1 on most work" while costing "40% less to run than Opus 5." The early independent record mostly supports that pitch, with two catches: at max effort the model talks enough to eat the discount, and on close inspection the capability jump over Fable 5.1 is incremental, not a new frontier.
That framing is not marketing spin cooked up in a vacuum. Within hours of release, the top-voted thread in r/ClaudeAI was titled "Opus 5.5 is 40% cheaper while being 30% faster than opus 5" — 775 upvotes of the community restating the official numbers as a buying argument. The same thread's top comment explains why the reaction is so intense: Opus 5 "couldn't even communicate properly," and people wasted "even more time and tokens" fighting it (u/Front_Raspberry_6488, 226 upvotes). This review sorts which of those feelings the evidence supports.
Every fact below was opened and checked on September 23, 2026, one day after release: the Anthropic launch page, the Artificial Analysis model page, and the original Reddit threads. The Claude Opus 5.5 model resource holds the prompt and evidence library; its review cards are the source-level record behind this article. No Opus 5.5 pricing or alternatives sibling page exists yet, so this review links none. At the end, I cover what the findings mean for browser-level workflows, where a model like this increasingly does its work.
Key takeaways
The one number is 40%, and it is a workload claim, not a price tag. Prices fell 20% ($4/$20) and cache reads fell 60% (to $0.20/M). The "40% less to run" is Anthropic's own testing of typical workloads — SonarSource independently measured a literal 40% fewer output tokens on Java generation, while Artificial Analysis measured $5.98 per index task at max effort.
Top of one independent leaderboard, at a cost. Artificial Analysis scores Opus 5.5 (max effort) at 58, #1 of 212 models, while ranking it #95 of 212 for verbosity: about 119k output tokens per index task versus about 27k for GPT-6 Astra max.
The communication problem that sank Opus 5 looks fixed. The loudest community complaint about Opus 5 was that coding was fine but talking to it was not. Day-one voices describe 5.5 as fast, natural and less padded, and SonarSource measured the comment share of generated Java dropping from 10.5% to 3.1%.
Less code, denser bugs. SonarSource's pre-release Java run found 27.5% less code and 42% fewer issues at a flat pass rate, but bug density per million lines rose from 576 to 644, with concurrency bugs up 44%. Review passes stay mandatory.
METR's read: incremental, not a leap. The predeployment evaluation found incremental gains over Fable 5.1 across five AI R&D tasks, with judgment-shaped weaknesses (foresight, feedback loops, research taste) still open. Buy 5.5 for the economics and the polish; escalate to Fable when you hit its ceiling.
What Claude Opus 5.5 actually is
Claude Opus 5.5 is the first model of Anthropic's new Claude 5.5 family, released September 22, 2026. The API model ID is claude-opus-5-5; the launch page lists availability on the Anthropic Claude Platform, AWS, Google Cloud and Microsoft Azure. The documented envelope is a 1M-token context window, up to 128K output tokens, adaptive thinking, and medium as the default effort — that last detail matters more than any benchmark row in this review.
| Question | Published snapshot (checked 2026-09-23) | What it does not prove |
|---|---|---|
| Release and model ID | 2026-09-22; claude-opus-5-5 | That every client exposes the same build or effort range |
| Input / output price | $4 / $20 per million tokens | Your completed-task cost after thinking and retries |
| Cache reads / writes | $0.20 / $5 per million tokens (Opus 5: $0.50 / $6.25) | A universal 40% saving on every workload |
| Context and output | 1M tokens in, up to 128K out | That every provider or plan exposes the full window |
| Effort | adaptive thinking; default medium | That lower effort keeps the same scores (curve not published per task) |
| Availability | Anthropic, AWS, Google Cloud, Azure | Consumer-plan access; not verified here |
| Independent snapshot | Artificial Analysis: 58, #1/212 (max effort) | Speed: AA lists latency as N/A for this model |
Anthropic also frames the release procedurally: it is the first release since the company "called for pacing the frontier," and it was tested before release by external evaluators including Frontier Design and METR. On Anthropic's own automated behavioral audit — about 2,000 scenarios — Opus 5.5 is "the strongest-performing model we've tested to date," and one containment-boundary evaluation showed roughly 85% fewer bypass attempts than Opus 5 or Mythos 5.1. Those are safety claims worth recording, and they cut both ways: a later section shows how safeguard interventions can also turn into zero scores on capability benchmarks.
The question this review answers is not "is it smart." A launch table from the vendor and a #1 independent snapshot already say yes. It is whether a model priced like a workhorse and pitched at flagship parity belongs in your daily workflow — versus keeping Opus 5 rates, paying Fable 5.1 prices, or moving mechanical work to cheaper models like GPT-5.6 Sol or MiMo-V2.6-Pro.
The one number that decides this review: 40%
Anthropic's pitch sentence, verified on the live release page, is worth reading precisely: Opus 5.5 "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." Two claims in one line — near-flagship capability, workhorse economics — and both are testable.
The price part is simple and public: input fell $5 → $4, output fell $25 → $20, cache reads fell $0.50 → $0.20 per million tokens. Cache reads matter most for agentic work because a 40-tool-call run re-reads its cached prefix 40 times; at $0.20 — 5% of the input price, what Artificial Analysis lists as a 95% cache discount — long sessions get structurally cheaper. That is the same lever that made Fable 5.1's cache-read cut the story of its release, now pushed further down.
The 40% part is where the fine print lives, and three independent measurements bracket it:
Collected: SonarSource's Java generation run (pre-release build, High effort) used 12.96M output tokens versus Opus 5's 21.71M — a literal 40% reduction — at a pass rate within one point (87.68% vs 88.6%). Less code, fewer tokens, same functional result. This is what the discount looks like when it works.
Spent: Artificial Analysis ran the Intelligence Index at max effort and measured about 119,000 output tokens per task — versus about 73k for Opus 5 max, about 78k for Fable 5.1 max, and about 27k for GPT-6 Astra max — landing at $5.98 per weighted index task and a verbosity rank of #95 of 212. This is what the discount looks like when the model is allowed to think out loud at full volume.
In between: Anthropic's own internal M&A analysis finished in 63 minutes versus Opus 5's 93 at "50% less to run" — vendor-tested, right at the claim.
The reconciliation is not a contradiction; it is the effort dial. Medium is the default for a reason: the 40% discount is real, and whether you collect it is mostly a function of which effort level your workflow reflexively uses. METR's capability read supports the same conclusion from the other side — the model is an incremental step over Fable 5.1, so the reason to switch is economics and polish, not new ceiling. If your pipeline pins max because "best model, best effort," you have re-priced yourself back to flagship territory.
What happened on real work, before the leaderboard
The most useful day-one evidence is not a benchmark row. It is three specific build stories, all from named users, all with their limits stated.
A 45-minute single-shot film. u/mshort3 gave Claude Code one sentence — "Lets make a ~70s riso film of your own choosing, free range" — inside an existing creative project, and reports Opus 5.5 produced a 78-second animated train journey in a single run of about 45 minutes (r/ClaudeAI, 637 upvotes). The repo is public, and the author credits the project's own skills and docs, not just the model. One run; no logs of the intermediate states.

Original thread: r/ClaudeAI 1wnkvys.
A one-minute video from one prompt — with a cost receipt. u/AzorAhai1TK asked for "an average day at work… with a twist" rendered as a creepy one-minute video, and reports the model coded and rendered it end-to-end at Extra-High effort in about 45 minutes, using roughly 4% of the weekly allowance on a $20/month plan (r/ClaudeAI). Self-reported numbers, but the quota figure is exactly the kind of datapoint a pricing debate needs: creative max-effort runs are not free, and they are not catastrophic either.
"Night and day" speed in Claude Code — with the caveat attached by its own author. u/Lazy_Assistance_1137 reports UI bugs that "used to take a while to track down" get spotted "in no time," calling the jump "night and day" — then immediately warns "this could just be the classic day-one honeymoon phase, and in two weeks we'll all be back to posting 'did they nerf it?' threads" (r/ClaudeAI, 273 upvotes). The top reply is one word deep: "It's crazy fast. I'm impressed." (u/capivara_de_pijama, 96 upvotes).

Original thread: r/ClaudeAI 1wnil7n.
Three stories, n=1 each, no logs, no token counts (except the self-reported 4%). What they establish: the model completes long, self-directed, multi-tool creative builds on the first day, repeatedly, in public, with artifacts you can open. What they do not establish: how often it does so. That is the gap benchmarks are supposed to fill — let us see how much they actually fill it.
The benchmarks, and how much to trust them
Anthropic's launch table, recorded as published on 2026-09-22. Different columns can come from different operators and leaderboards (the GPT-6 Astra and GPT-5.6 Sol figures are identified by Anthropic as coming from OpenAI reports), so read rows, not column arithmetic:
| Domain / benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Agentic coding: Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| Agentic coding: FrontierCode v1.1 | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| Agentic coding: CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| Knowledge work: GDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| Business workflows: AutomationBench | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Reasoning: Humanity's Last Exam (tools) | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Agentic science: Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| Computer use: OSWorld 2.0 (partial) | 81.8% | 80.7% | 74.0% | — | — |
| Chart recognition: Chartography (tools) | 89.0% | 88.4% | 83.4% | — | — |
Now the cold water, in three parts.
The error bars are published, and they are wide. Terminal-Bench 4.0 carries a ±2.6-point standard error for Opus 5.5; Terminal-Bench-Science carries ±3.5–5 points per model. The Opus 5 → Opus 5.5 jump on TB-Science (29.0% → 58.7%) survives any error bar; the gap to GPT-6 Astra (64.6%) on the same row also survives it — in Astra's favor. Anthropic itself cautions on the page that at current capability levels, benchmark score differences may not reliably represent real-world gaps. That sentence, from the vendor, is the most honest line in any launch table this cycle.
The comparison conditions are not uniform. Terminal-Bench rows use the Claude Code harness; AutomationBench was run and reported by Zapier, without fallback models, so safety interventions counted as failures; GDPval-AA v2.1's independent scoring process is not detailed on the page. Safeguarded interventions can zero out a scored task — a policy outcome recorded as a capability miss. None of this makes the table wrong; it makes it a set of differently-conditioned snapshots, not a controlled race.
The independent record agrees on shape, not on margins. Artificial Analysis's own index (v4.3.2, ten evaluations) puts Opus 5.5 max at 58, #1 of 212, leading six evaluations outright (Humanity's Last Exam 61.4%, SciCode 66.9%, GDPval-AA 1846, AA-Briefcase 1822 — 143 Elo above Fable 5.1) and tying GPT-6 Astra xhigh on Terminal-Bench 4.0 at 59.6%. SonarSource's controlled Java test is the sharpest single datapoint: pass rate within a point of Opus 5 while generating 27.5% less code — and bug density per line up, covered below. METR's predeployment evaluation, run over 10 business days on five AI R&D tasks, concludes the gains over Fable 5.1 are "incremental… rather than a discontinuous leap," with no major progress on foresight, feedback-loop building or research judgment.
What survives the dousing: the before/after gap over Opus 5 is real and large (TB-Science 29.0% → 58.7% is not a rounding trick), the #1 independent snapshot is corroborated by a second leaderboard-style source, and the vendor's cost story has one fully independent replication (SonarSource's token counts). What does not survive: any claim that Opus 5.5 cleanly "beats" GPT-6 Astra. On Anthropic's own table Opus 5.5 leads Terminal-Bench 4.0 while Astra leads Terminal-Bench-Science; Artificial Analysis has the two tied at 59.6% on its independent TB-4.0 run; and our GPT-6 Astra review covers that model's own trade-offs.
Where Opus 5.5 is genuinely good
It finally talks like a workmate
The strength no benchmark row captures: Opus 5 was a strong coding model people hated conversing with, and day-one voices say 5.5 fixes the conversation itself. The r/ClaudeAI anchor thread's top comments are a post-mortem of that pain — "Opus is fantastic, just until you have to talk to it" (u/Neon_Camouflage, 61 upvotes) — paired with relief: "Heard Opus 5.5 is less verbose" (u/No-Temperature6597). SonarSource's analyzer puts a number on the behavior change: the comment share of generated Java fell from 10.5% of lines to 3.1%, with code-smell density down 21%. And the trained-philosopher vibe check (medium effort, incognito) describes "a model with polish and panache… It will happily cite back Theophrastus and Austin to you, but can't break the habit of wanting to get its hands dirty" (u/Wickywire). If you write or plan for a living — not just compile — this is the upgrade that matters more than any coding score.
Agentic follow-through, at Opus pricing
This is the official table's center of gravity, and it holds up: 66.4% Terminal-Bench 4.0 and 58.7% TB-Science (both ± their published error bands) with the Claude Code harness; 57.8% CursorBench 4.0; OSWorld 2.0 at 81.8% partial for computer-use work. The internal examples are unusually concrete for vendor evidence: on a quarterly-earnings task, 16 of 18 reports passed a grader that checked every figure and citation and failed both Fable 5.1's and Opus 5's attempts; on a fictional M&A analysis, 63 minutes versus Opus 5's 93 at half the cost. Vendor-tested, but exactly the shape of work — long, multi-file, verification-heavy — that the terminal numbers measure. For agentic research and deep-work workloads, this profile at $4/$20 is the strongest pitch of the release.
Medium-effort economics are the quiet headline
Default medium plus $0.20 cache reads changes unit economics for long sessions in a way headline prices do not. Artificial Analysis placed four of the five effort levels on the intelligence-vs-cost Pareto frontier — the exception being the level nobody should reflex-choose — and its $5.98/task figure belongs to max, not medium. CodeRabbit's pipeline test found the Standard configuration (lower effort) beat its own Max configuration on the 80-pattern set with better actionable precision and fewer comments, while Max pulled ahead only on the 13 hardest Signal cases. The pattern across all three independent sources is consistent: this model's value function rewards deliberate effort selection, and punishes reflex. For teams currently paying Opus 5 rates — or legacy Opus 4.8 rates — the switch is close to free upside.
Where it bites
Max effort burns the discount
The same dial that makes medium cheap makes max expensive: about 119k output tokens per AA index task (4.4× GPT-6 Astra max, 1.6× Fable 5.1 max), $5.98 per weighted task, verbosity ranked #95 of 212, cumulative 260M tokens across the evaluation. CodeRabbit saw the same shape in production-shaped work: token usage up 49–60% versus its baseline, with Max adding 57.6% (OSS set) to 60.1% (Signal set) over baseline for precision it often did not need. If your harness or habit pins max effort, budget frontier money and see also Gemini 3.8 Flash's review for what a deliberately frugal alternative looks like — and MiMo-V2.6-Pro for the cheapest measured cost per index task in the recent cohort.
Less code, denser bugs
SonarSource's numbers deserve their uncomfortable framing: total issues fell 42%, but bug density per million lines rose 576 → 644, and concurrency/threading bugs — the class that surfaces in production, not in review — rose 44% to become the largest category (295/mLOC). Every BLOCKER-severity density fell, which is genuinely reassuring, but the shape is clear: 5.5 writes tighter code that is not uniformly safer code. Two boundaries: the run used a pre-release build (SonarSource will re-run at general availability), and CodeRabbit's parallel finding — 11 new patterns found, 9 baseline patterns missed on the OSS set — says the same thing from the review side. Keep the review pass; move it earlier in the pipeline, don't delete it.
Incremental over Fable, and the unknowns that stay unknown
METR's evaluation is the least flattering independent read, and it is worth reading twice: gains over Fable 5.1 across all five AI R&D tasks, no evidence of full automation, and no major progress on the judgment-shaped capabilities — foresight, prediction, building feedback loops, research taste. Add the genuinely unknown: no verified latency exists (Artificial Analysis lists speed as N/A for this model, so this review quotes no TTFT number), consumer-plan availability is unverified, and safeguarded runs can zero-score tasks, which means your harness's safety settings are part of your measured capability. A one-day-old model ships with all of that attached. None of it is disqualifying; all of it argues for a measured pilot rather than a fleet-wide flip.
What people actually said
Day-one conversation splits along two axes, not one.
Axis one: finally fast — and finally talkable. The speed thread and its replies are unambiguous delight; the philosophy interview is the qualitative version of the same thing, describing polish and an eagerness to "get its hands dirty." A user rebuilding workflows around it sums up the constituency: people who liked Opus 5's output and hated the conversation.

Original thread: r/ClaudeAI 1wnkgie.
Axis two: trust, but verify — the Opus 5 scar tissue. The anchor thread's skepticism is specific: people were burned by a model that benchmarked well and communicated badly, and they are not accepting launch-day numbers as closure. One commenter reads the design intent charitably — "Fable 5 & Opus 5 were designed & tested together with the intent being that you talk to Fable and it delegates execution to Opus agents" (u/canyonero7) — which raises the real question about 5.5: it is priced as the execution model but talked up as the conversation model, and both roles in one model is exactly what day-one hype cannot confirm.

Original thread: r/ClaudeAI 1wnf7sb.
Where this review lands: with the enthusiasts on the switch decision — the economics are too one-sided to keep new work on Opus 5 — and with the skeptics on the process. The honeymoon-phase author is right that two weeks of logs will tell you more than any table above; he is also right that right now, it feels like a real jump.
The verdict: switch your Opus work, keep the discipline
| Workload shape | Verdict | Why | Main caveat |
|---|---|---|---|
| Default coding-agent work currently on Opus 5 or 4.8 | Switch now | 20% lower prices, 60% cheaper cache reads, communication fix, TB-Science 29.0% → 58.7% | Pin the model ID; re-measure after the GA re-runs land |
| Hardest frontier and long-horizon research | Escalate to Fable 5.1 | METR: incremental over Fable; official pitch itself is "level of Fable on most work" | Fable's own costs and quirks are a separate decision |
| Unattended long runs, batch pipelines | Good candidate at medium | $0.20 cache reads + default medium effort; AA Pareto frontier at 4 of 5 levels | No verified latency — measure TTFT yourself before betting SLAs |
| Interactive chat products | Measure first | Unknown latency; verbosity at high effort (#95/212) directly hits UX and bill | AA lists speed as N/A; no number quoted here |
| Java or large code generation without review | Use, with review moved earlier | Pass rate flat, output leaner, BLOCKER densities down | Bug density up 576 → 644/mLOC; concurrency +44% |
| Cost-floor automation | Compare cheaper models | $4/$20 is cheap for Opus-class, not cheap in absolute terms | MiMo at $0.13/task or Gemini 3.8 Flash may fit mechanical work |
Two time-sensitive notes. First, every independent number in this article is a day-one snapshot: SonarSource tested a pre-release build and says it will re-run at general availability; Artificial Analysis rankings move daily. Second, the family ladder is still filling in — Sonnet 5 and Haiku 4.5 sit below this release, and Anthropic's "Claude 5.5 family" wording says more members are coming. Plan model routing with a pin, not a habit.
A practical route: run it where the work happens
Everything above points at one pattern: the economics reward long, tool-using runs at deliberate effort, the failure mode is unsupervised volume, and the open questions — latency, picker availability, stability over weeks — are all things you learn by running one bounded task, not by reading one more table. That is a browser-shaped job: multi-page research, extraction and automation where an agentic browser holds the session's pages and context while the model works, with Tabbit Browser built around exactly that loop.
The honest boundary, same as our other model reviews: this review did not verify whether Opus 5.5 appears in Tabbit Browser's live model picker, and no hands-on Tabbit run is claimed. Check your account's model selector, pin the model ID, give it one reversible task — a multi-file change with an existing test, or a research packet with required citations — and log cost per accepted result against Opus 5 before you switch anything you care about.
Sources
Anthropic: Introducing Claude Opus 5.5 — release facts, pricing, benchmark table, safety section (opened 2026-09-23)
Artificial Analysis: Claude Opus 5.5 and launch article — index 58 (#1/212), $5.98/task, verbosity measurements
METR: Predeployment evaluation of Claude Opus 5.5 — incremental-over-Fable read
SonarSource: Evaluating Claude Opus 5.5 on Java code generation — pass rate, token and bug-density numbers (pre-release build)
CodeRabbit: Opus 5.5 model review — recall/precision and token-usage trade-offs
FAQ
Is Claude Opus 5.5 worth it?
Yes as your default Opus-class model. It aims at Fable 5.1-level output on most work at $4 per million input and $20 output tokens, 20% below Opus 5, with cache reads at $0.20. Keep Fable 5.1 for the hardest frontier tasks and avoid reflex max-effort runs, where measured verbosity can erase the savings.
Is Claude Opus 5.5 cheaper than Opus 5?
List prices are 20% lower for input ($4 vs $5) and output ($20 vs $25), and cache reads drop from $0.50 to $0.20 per million tokens. Anthropic claims about 40% lower total cost for typical workloads based on its own testing. Your actual saving depends on effort level and token mix, so measure a representative task.
Is Opus 5.5 better than Fable 5.1?
Anthropic says Opus 5.5 performs at the level of Fable 5.1 on most work while costing less to run. Independent checks add nuance: METR calls the AI R&D gain incremental over Fable 5.1, while Artificial Analysis scores Opus 5.5 at 58, the top of its 212-model snapshot. Treat it as near-parity at lower cost, not a clean win.
Is Claude Opus 5.5 good for coding?
The launch table reports 66.4% on Terminal-Bench 4.0 with a plus or minus 2.6 point standard error and 57.8% on CursorBench 4.0. SonarSource's Java test found a broadly similar pass rate to Opus 5 (87.68% vs 88.6%) with less code but higher bug density, so keep a review pass instead of trusting output blind.
Why does Claude Opus 5.5 use so many tokens?
At max effort, Artificial Analysis measured roughly 119,000 output tokens per Intelligence Index task, more than four times GPT-6 Astra's roughly 27,000, ranking Opus 5.5 95th of 212 for verbosity. The default effort is medium, and SonarSource measured 40% fewer output tokens than Opus 5 on Java generation, so the burn is effort- and workload-specific.
When was Claude Opus 5.5 released, and where is it available?
Anthropic released it on September 22, 2026, with API model ID claude-opus-5-5, a 1M-token context window and up to 128K output tokens. The launch page lists the Anthropic Claude Platform, AWS, Google Cloud and Microsoft Azure. Consumer-plan availability was not verified for this review.
Can I use Claude Opus 5.5 in Tabbit Browser?
This review did not verify whether Opus 5.5 appears in Tabbit's live model picker, so do not plan around it. Check your own account's model selector and run one reversible task before committing a workflow to it.