Claude Opus 4.7 is best understood as a focused upgrade to Opus 4.6: stronger software engineering, tool use, computer interaction and vision, with a visible regression on the published BrowseComp result. It is not a universal leaderboard winner, and it is no longer the current Opus destination for a new build.
That is the decision anchor as of September 20, 2026. Anthropic's live model page lists claude-opus-4-7 as Active (legacy), keeps the standard API price at $5 per million input tokens and $25 per million output tokens, and recommends considering Opus 5. The practical question is therefore two questions: does 4.7 fit this workload, and is the 4.7-to-5 migration cost worth avoiding today? Tabbit can be a browser-level place to test a reversible workflow, but this draft contains no signed-in Tabbit test.
Key takeaways
Opus 4.7's strongest public story is difficult coding, tool orchestration, desktop interaction and chart or interface understanding.
The live catalog lists 1M context, 128K maximum output, adaptive thinking,
highdefault effort and a January 2026 reliable-knowledge cutoff.The reported table is mixed: SWE-bench Verified rises from 80.8% to 87.6%, while BrowseComp falls from 83.7% to 79.3%.
The newer tokenizer produces approximately 30% more tokens for the same text, according to Anthropic's pricing documentation. List price is unchanged; task cost need not be.
Opus 4.7 remains available through Anthropic's API and listed cloud routes, but a new production integration should compare Opus 5 and record a rollback plan.
Claude Opus 4.7 at a glance
The Claude Opus 4.7 model resource contains the source-level prompt and review material. This overview keeps the version, access and lifecycle decision together.
| Question | Current catalog snapshot checked 2026-09-20 | Decision boundary |
|---|---|---|
| API model ID | claude-opus-4-7 | Pin the exact ID in logs; provider aliases may differ. |
| Released / status | April 16, 2026 / Active (legacy) | Available does not mean current family endpoint. |
| Context / maximum output | 1M tokens / 128K tokens | Catalog ceilings; clients and providers can expose less. |
| Input / output | Text and images / text | Tool permissions come from the route and harness. |
| Thinking / default effort | Adaptive / high | Compare effort, tokens and completion quality together. |
| Standard API price | $5 input / $25 output per MTok | Cache, Batch, cloud and consumer-plan terms are separate. |
| Providers listed | Claude API, Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS | Region, quota and feature availability still need checking. |
| Retirement boundary | Not sooner than April 16, 2027 | “Not sooner” is not a guaranteed final date. |
The prompt resources and review resources are better for source-by-source instructions and methodology. Do not read a model-resource count as a production success rate.
What changed from Opus 4.6?
Anthropic describes Opus 4.7 as a meaningful improvement on advanced software engineering, especially difficult tasks that require a model to plan, use tools and verify its work. The release also raises the image ceiling to roughly 3.75 megapixels, with a 2,576-pixel long edge. That matters for dense diagrams and interfaces, not just for prettier image descriptions.
The launch announcement reports the following practical direction:
Harder coding: the published SWE-bench and Terminal-Bench results improve over 4.6.
More reliable tool loops: MCP-Atlas and OSWorld-Verified move upward, which is relevant to agents that must act rather than only explain.
Sharper visual reasoning: CharXiv improves substantially with and without tools.
Stricter deployment boundaries: Anthropic says Opus 4.7 is less capable than Mythos Preview for cyber tasks and ships automatic safeguards for prohibited or high-risk requests.
The change is not “4.7 is smarter at everything.” Anthropic's release and the Vellum analysis both preserve an important caveat: the BrowseComp result is lower than 4.6. A research agent that spends most of its time searching and synthesizing pages deserves a separate evaluation.
The benchmark table is a workload map, not a ranking
These figures come from Anthropic's published results and Vellum's reading of the same evidence. They use different benchmark harnesses and do not predict your application's success rate.
| Benchmark | Opus 4.6 | Opus 4.7 | What the movement suggests |
|---|---|---|---|
| SWE-bench Verified | 80.8% | 87.6% | Stronger issue-resolution signal on this benchmark. |
| SWE-bench Pro | 53.4% | 64.3% | Larger gain on the harder multi-language variant. |
| Terminal-Bench 2.0 | 65.4% | 69.4% | Better command-line task performance in the reported run. |
| MCP-Atlas | 75.8% | 77.3% | Incremental improvement in multi-turn tool use. |
| OSWorld-Verified | 72.7% | 78.0% | Better computer-use result in the published table. |
| BrowseComp | 83.7% | 79.3% | A real regression for web-research-style tasks. |
| CharXiv, no tools / tools | 69.1% / 84.7% | 82.1% / 91.0% | The clearest visual reasoning improvement in the table. |
The release page does not provide every task input, tool schema, repetition count or confidence interval. Vellum also notes that the numbers should be read by task. The responsible conclusion is conditional: try 4.7 for code and tool-heavy work, but do not remove retrieval checks because a coding score went up.
Access, pricing and the token boundary
The current pricing documentation lists Opus 4.7 at $5 per million input tokens and $25 per million output tokens. It also lists $6.25 for a five-minute cache write, $10 for a one-hour cache write, $0.50 for cache hits and refreshes, and a 50% Batch API discount on input and output. These are API rules, not a Claude.ai subscription quote.
Anthropic also says Claude 4.7 and later use a newer tokenizer that produces approximately 30% more tokens for the same text, depending on content and workload shape. This is the number to put beside the unchanged list price: a long prompt or agent trace can still consume more billable tokens after a model upgrade.
| Route | What is documented | What you must verify |
|---|---|---|
| Claude API | claude-opus-4-7, standard token pricing, adaptive thinking | Organization limits, prompt caching, data-residency route and actual token counts. |
| Bedrock / Google Cloud / Microsoft Foundry | Provider model IDs are listed | Region, endpoint premium, quota, tools and provider billing. |
| Claude products and Claude Code | Product access is separate from API access | Plan, usage pool, selector and effort controls in the live account. |
| Tabbit Browser | Browser workflow may provide a practical client surface | Selector visibility, effective context, latency, cost and tool permissions; not verified here. |
Do not combine API unit price with a subscription seat or with Tabbit's product availability. For the adjacent family, the Claude Opus 4.8 overview covers the next release's own lifecycle and evidence, while the Claude Sonnet 5 overview covers the lower-cost route.
Unknown risks worth testing before adoption
The public evidence points to four practical risks:
Effort can hide cost: adaptive thinking and higher effort can improve hard tasks but increase thinking tokens, waiting and rate-limit use.
Ambiguity may be a workflow problem: a Reddit user reported better instruction following but less reliable interpretation of underspecified prompts than their 4.6 baseline. Treat that as a prompt and harness observation, not a universal regression.
Dense answers can slow review: another Reddit user found Opus 4.7's planning output hard to parse and used a separate session to translate it. If your team makes decisions from model plans, measure review time as well as code quality.
Confident mistakes still need a second pass: the community also reports stronger code and more thorough reviews alongside lower trust when answers sound right but are wrong. Require tests, citations or an independent reviewer.
For web research, the published BrowseComp fall is enough reason to preserve source citations and contradiction checks. For irreversible browser or terminal actions, require confirmation regardless of benchmark performance. The agentic browser guide explains why a model's tool capability and a browser's permission boundary are separate decisions.
Community signal: useful disagreement, not a score
The Reddit evidence is unusually concrete but still personal. One heavy Claude Code user wrote that 4.7 followed instructions and produced better code, while being “less consistent than 4.6”; another thread described the output as difficult to understand during a pathfinding-plan review. A later side-by-side code review, dated August 27, reported 6 of 8 Opus 4.7 claims confirmed after checking the code, with one inflated and one wrong.
These reports agree on the operational lesson, not on a universal verdict: effort, prompt specificity, task design and review harness matter. The original discussions are the heavy-use comparison, the readability complaint and the dated code-review comparison. They are not controlled tests, and this draft has no approved screenshot artifacts.
A fourth discussion, the “medium” effort thread, shows the same split: one commenter saw details that 4.6 missed, while another described an unreasonable formatting decision. That disagreement is useful evidence about workload and context sensitivity, not a model-wide score.
Migration: 4.7 is usable, but new builds should look ahead
The live Opus 5 migration guide says migration from 4.7 or earlier includes more than changing the model string. Anthropic calls out older sampling and prefill behavior, manual thinking, the newer tokenizer and response handling. A tool loop must preserve thinking blocks exactly when the target route emits them; code that assumes content[0] is always text can break.
Use this short self-check before switching:
Pin
claude-opus-4-7and the candidate target in the same harness.Record effort,
max_tokens, input/output tokens, cache hits, tool calls, latency and retries.Run one coding task, one tool workflow, one visual task and one web-research task.
Keep tests, citations and a human approval step in the acceptance criteria.
If the target changes thinking defaults or response blocks, test the adapter and rollback path before changing the production alias.
| Your priority | First route to evaluate | Why |
|---|---|---|
| Difficult repository fixes and tool orchestration | Opus 4.7 alongside Opus 5 | 4.7's public gains are clearest here, but lifecycle makes the newer route relevant. |
| Dense screenshots, diagrams or UI agents | Opus 4.7 with a visual acceptance set | CharXiv and OSWorld move up; actual UI permissions still matter. |
| Research-heavy browser agent | 4.7 versus a current route with retrieval checks | BrowseComp is lower than 4.6; do not infer research quality from coding results. |
| Lower-cost daily coding | Sonnet 5 | Compare the Sonnet 5 page on its own price and effort evidence. |
| New long-horizon, high-stakes build | Opus 5 or Fable 5.1 pilot | Review the Fable 5.1 analysis and test the current route before pinning a legacy model. |
A practical Tabbit boundary and next step
Tabbit Browser is relevant when the question is not only “what can the API generate?” but “can I inspect a live page, keep sources visible and review the resulting action?” The browser automation guide, AI browser guide and Tabbit Browser overview explain that client-level workflow.
This draft did not run a signed-in Opus 4.7 task. It therefore does not claim that the model appears in a selector, that the browser exposes 1M context, or that latency and pricing match Anthropic's API. If the model is visible in your account, choose a reversible research task, record the model label and effort, save the sources, and judge the final answer against a written acceptance test.
Verdict
Claude Opus 4.7 is a meaningful Opus 4.6 upgrade for hard coding, tool use and visual reasoning. Its reported BrowseComp regression, newer-tokenizer cost boundary and mixed community reports rule out a blanket “upgrade everything” recommendation. It remains a reasonable controlled pilot while its API route is active, but a new production build should compare Opus 5 now and preserve a migration path.
Sources
Primary sources are Anthropic's launch announcement, Opus 4.7 model page, pricing documentation, model deprecations, Opus 5 migration guide and Opus 4.7 system card. The independent interpretation is Vellum's benchmark explanation. Community evidence is linked in the disagreement section.
FAQ
What is Claude Opus 4.7?
Claude Opus 4.7 is Anthropic's April 2026 Opus release, aimed at difficult software engineering, tool-heavy agents and higher-resolution vision. The current model page lists claude-opus-4-7 with a 1M-token context window, 128K maximum output and adaptive thinking.
What changed from Claude Opus 4.6?
Anthropic reports gains in coding, tool use, computer interaction and visual reasoning, while BrowseComp is lower than the 4.6 result in the published table. The tokenizer and effort settings also affect completed-task cost, so the upgrade should be tested on a fixed workload.
How much does Claude Opus 4.7 cost?
The current Claude Platform page lists $5 per million input tokens and $25 per million output tokens, with separate cache and Batch rules. API rates are not the same as Claude subscriptions, cloud-provider invoices or Tabbit availability.
Is Claude Opus 4.7 still available?
Yes, the live model page labels it Active (legacy) and says retirement is not sooner than April 16, 2027. Anthropic also recommends considering Claude Opus 5, so a new integration should compare the newer route before committing.
Should I migrate from Opus 4.7 to Opus 5?
Run a fixed acceptance set first, then review thinking defaults, max_tokens, response-block handling, tool loops, tokenizer changes and safety fallbacks. A model-ID swap alone is not enough for a production migration.
Can I test Claude Opus 4.7 in Tabbit?
This draft did not run a signed-in Tabbit Opus 4.7 task, so it does not claim selector visibility, effective context, latency, cost or tool access. Check the live selector and use one reversible task with a written acceptance test.