Release notes for the GPT-5.6 series, including official results for Sol on Agents’ Last Exam, the Artificial Analysis Intelligence Index, coding, knowledge work, and safety evaluations, as well as model positioning, ultra mode, and pricing information.
The following is the visible body text extracted during this visit. It includes page navigation, machine translation, advertising, comments, and other page elements; verify against the original link before citing it.
Skip to main content Research Products Business Developers Company Foundation (Opens in a new window) Log in Try ChatGPT (Opens in a new window) GPT-5.6: Frontier Intelligence That Scales Flexibly to Ambitious Goals | OpenAI
July 9, 2026
Products Release GPT‑5.6: Frontier Intelligence That Scales Flexibly to Ambitious Goals
More intelligence in every token, better performance per dollar, and stronger capabilities on demand for the most challenging tasks
Listen to article 23:11 Share Efficient by default, powerful when needed A leap forward in design capabilities End-to-end knowledge work Pushing the frontier of cybersecurity and science GPT-5.6 accelerates OpenAI Safety and security that scale with model capability Availability and pricing
00:00
Updated July 30, 2026: OpenAI has reduced the price of GPT‑5.6 Luna by 80% and GPT‑5.6 Terra by 20%. Learn more here.
After the end of the limited preview, we are officially launching the GPT‑5.6 model series, including the new flagship model Sol, the balanced Terra for everyday work, and the highly cost-effective Luna.
GPT‑5.6 Sol sets a new standard for intelligence and efficiency, delivering frontier-level results in coding, knowledge work, cybersecurity, and science while outperforming previous-generation and other frontier models with fewer tokens and lower estimated costs. This translates into better performance per unit cost: it can complete more tasks on the same budget, or achieve the same result at a lower total cost. We are also introducing a new way to accelerate the heaviest workloads: ultra is our highest-performance setting, coordinating multiple agents across parallel workflows to complete complex tasks faster. Stronger computer-use capabilities and design judgment make GPT‑5.6 Sol our best assistant yet for collaborative experiences, able to inspect, polish, and deliver results that are ready to use.
Through training GPT‑5.6, we made every token produce more useful results. On Agents’ Last Exam (Opens in a new window) , a long-horizon professional-workflow evaluation spanning 55 fields, GPT‑5.6 Sol set a new high of 53.6, leading Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning effort, it led Fable 5 by 11.4 points at roughly one-quarter of the estimated cost. This efficiency advantage also appears in smaller models, which is essential to making intelligence more widely available and affordable: GPT‑5.6 Terra and GPT‑5.6 Luna outperformed Fable 5 at roughly one-sixteenth of its cost. On the Artificial Analysis Intelligence Index (Opens in a new window) , a broad standard covering metrics such as agentic work, coding, scientific reasoning, and general capabilities, GPT‑5.6 Sol with max reasoning effort came within 1 point of Fable 5, while completing tasks 61% faster at roughly half the estimated cost.
Agents' Last Exam Artificial Analysis Intelligence Index v4.1 Agents' Last Exam Cost Latency Output tokens Cost $0 $1,000 $2,000 $3,000 $4,000 30% 40% 50% Score API cost (USD) GPT-5.6 Sol GPT-5.6 Terra GPT-5.6 Luna GPT-5.5 Claude Fable 5 Claude Opus 4.8 Gemini 3.1 Pro Preview
Agents’ Last Exam (Opens in a new window) : Long-horizon agentic workflows across professional disciplines.
GPT‑5.6 comes with our most robust safeguards yet, designed to resist deliberate and evolving abuse without unduly restricting legitimate work. Before the broad launch, we combined human red teaming with large-scale automated testing in the largest evaluation of a model and its safeguards in our history. During the preview, we worked closely with expert organizations and trusted partners to stress-test and further strengthen the defenses ahead of a broader release. The final system uses a layered protection architecture that combines model-level defenses with real-time checks and monitoring, calibrating access according to trust and risk levels.
Efficient by default, powerful when needed
GPT‑5.6 Sol is our best coding model yet. On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning effort set a new SOTA at 80, 2.8 points above Fable 5, while using more than half as many output tokens, taking more than half as much time, and costing about one-third less. This advantage runs across the product line: Terra performed slightly better than Fable 5, while Luna outperformed Opus 4.8; both did so in roughly one-third the time, with half the output tokens, and at about one-quarter of the competitor’s estimated cost. It also set new frontier results on Terminal-Bench 2.1 and DeepSWE, benchmarks focused on complex command-line workflows and long-horizon engineering in real codebases.
Artificial Analysis Coding Index Terminal-Bench 2.1 DeepSWE v1.1
Artificial Analysis Coding Agent Index: An independent index that evaluates coding agents on feature implementation, terminal operations, and performance in real codebases.
GPT‑5.6 can write and run lightweight programs, coordinate tools, process intermediate results, monitor progress, and choose its next action during a task. This allows tool-heavy tasks to proceed with fewer tokens, fewer model interactions, and less human intervention. Developers do not need to script every step or send every tool response back to the model. The programmable tool-calling (Opens in a new window) feature in the Responses API can filter large amounts of intermediate data, retain only what matters, and dynamically adjust the workflow while it runs.
For problems worth more time and compute, GPT‑5.6 can go beyond this efficient default mode. The max option gives it more time than xhigh to reason, explore alternatives, validate, and continually revise its approach. ultra goes further by coordinating four agents in parallel by default. By using more tokens, ultra can produce stronger results on difficult tasks while returning them faster. The chart below compares ultra’s default four-agent configuration with a single-agent baseline on BrowseComp, SEC-Bench Pro, and Terminal-Bench 2.1; BrowseComp and SEC-Bench Pro also show a 16-agent configuration. Across all three evaluations, adding parallel agents shifts the score-latency frontier up and to the left, achieving better results in less time. In the API, developers can use the multi-agent (Opens in a new window) beta feature in the Responses API to build an experience similar to ultra. 4, 5, 6
BrowseComp (multi-agent) SEC-Bench Pro (multi-agent) Terminal-Bench 2.1 (multi-agent) 1 of 11 “ GPT‑5.6 is one of the most capable models we’ve tested on CursorBench, with strong early evaluation results. The gains in task persistence, intelligence, and overall efficiency are an exciting step for developers. We look forward to bringing this model to our Cursor users.” —Oskar Schulz, President of Cursor Cursor Qodo Notion Cognition Rogo Ramp Shopify Cisco Clio Balyasny Asset Management Basis A leap forward in design capabilities
GPT‑5.6 represents a major improvement in design aesthetics and judgment. With only high-level guidance, GPT‑5.6 can create interfaces that are beautiful, ergonomic, and fully functional. With stronger computer-use capabilities, it is no longer limited to generating underlying code or content; it can inspect and optimize the rendered result, proactively catching visual and functional issues and polishing details before delivering the final work.
Sailing game Tiny Voids game Museum website Clockwork Village game Interior design presentation
Prompt: Can you implement a 3d sailing game for me? For anything that needs bitmaps/textures/sprites (or if it helps to have a mockup reference for any 3d models you build) feel free to use imagegen.
In ChatGPT Work, GPT‑5.6’s frontend development capabilities can likewise turn natural-language requests into polished, interactive explanations and visualizations.
Interactive spirograph Interactive wave interference Interactive GPT tokenizer explainer
00:00
Prompt: Create an interactive spirograph to explain how it works.
End-to-end knowledge work
GPT‑5.6 delivers better results for professional tasks. It can extract messy background context from your documents and everyday workflows such as Slack, Notion, Microsoft 365, and Google Drive, then turn it into expert-level, shareable work.
GPT‑5.6’s knowledge-work advantage appears across evaluations covering long-horizon professional analysis, web browsing, tool use, and computer use. GPT‑5.6 Sol set new SOTA results of 92.2% on BrowseComp and 62.6% on OSWorld 2.0; on OSWorld, it outperformed Opus 4.8 while using 85% fewer output tokens. Here too, the entire GPT‑5.6 series improves performance per unit cost. Luna nearly reached GPT‑5.5’s best performance at less than half its estimated cost, while Terra outperformed GPT‑5.5 at lower cost.
BrowseComp GDPval-AA v2 OSWorld 2.0 AutomationBench
BrowseComp: GPT‑5.6 Sol set a new SOTA on BrowseComp, which includes agentic web-browsing tasks.
GPT‑5.6 improves the quality of presentations, documents, and spreadsheets, making its output more polished and accurate. It can create fully editable presentations from scratch, turning prompts and source materials into coherent visual narratives with excellent typography, hierarchy, and design.
Loading...
– / –
The improvement is especially clear when following templates and reference presentations. GPT‑5.6 can infer a presentation’s design system — including layout, fonts, spacing, colors, and recurring content patterns (including rules built into the slide master) — and apply those conventions consistently to new material. In this example, when asked to update figures based on a reference file, GPT‑5.5’s output omitted key components from the slide master, while GPT‑5.6 followed the reference structure more faithfully.
Reference file GPT‑5.5 output
GPT‑5.5 omitted key components from the slide master
GPT‑5.6 output
GPT‑5.6 also generates more visually refined documents and spreadsheets. It follows complex reference formatting more accurately, which is essential for repetitive knowledge-work tasks. It handles equations and financial models with greater precision and makes better use of typography, spacing, hierarchy, and page or worksheet layout.
Stock research document Leveraged buyout model Pinecrest Research Partners | Blossom Co. (BLSM) | Initiation
Please see important disclosures at the end of this report
1
EQUITY RESEARCH | Consumer Discretionary
Specialty Retail & Digital Commerce
8 July 2026
Pinecrest Research Partners LLC
Blossom Co. (NASDAQ: BLSM)
Recurring mix and delivery density create an earnings inflection
initiate at Overweight
Rating:
OVERWEIGHT
(initiation)
|
Price Target
$34.00
|
Last Close (7 July 2026)
$27.40
Implied Upside to Price Target: +24.1%
Market data
Market capitalization
$2.47 billion
Enterprise value
$2.31 billion
Net cash / (debt)
$153 million
Diluted shares outstanding
90.0 million
Free float
~88%
day daily volume
1.20 million shares
week range
$18.20
$31.60
Dividend yield
N/A (no dividend)
Trailing 12M ROE
12.1%
Fiscal year end
December 31
Listing / index
NASDAQ / Russell 2000
Reporting currency
USD
price series through 7 July 2026. Russell 2000 rebased to 100 at 8 July 2025.
Executive summary
We initiate coverage of Blossom Co. ("Blossom" or "BLSM") with an Overweight rating and a $34 price
enabled premium floral, gifting,
and subscription platform that combines proprieta ry personalization, nine regional preparation hubs, and
driven online florist into
frequency gifting platform, with recurring membership and enterprise services lift ing revenue
visibility, customer lifetime value, and fulfillment density.
Our investment thesis rests on three points:
•
(1) Revenue quality is improving.
platform revenue
represented 35.8% of FY25 sales; we model that mix reaching 44.3% by FY28. These streams carry
off
occasio n purchases. Paid subscribers rise from 1.15 million to 2.18 million over the same period,
recognition programs, loyalty integrations, and an
expanding fulfillment API.
BLSM 12-month price evolution vs Russell 2000
BLSM (price, $, left)
Russell 2000 (rebased = 100, right)
$18
$22
$26
$30
96
100
104
108
112
Jul 25
Sep 25
Nov 25
Jan 26
Mar 26
May 26
Jul 26
52-wk high $31.60
Last $27.40
BLSM price ($)
Russell 2000 (rebased = 100) 1 / 11
Early GPT‑5.6 testers found improvements in knowledge-work output across fields.
1 of 9 “ GPT‑5.6 is highly efficient in the long and complex workflows required to build production-grade applications. As one of the models currently used by Lovable, it helps users complete tasks with about 25% fewer steps and 35–48% fewer tool calls than the previous-generation model, while increasing project success rates by 15% and reducing hangs during runs. This is a meaningful improvement for anyone who wants to turn an idea into a working application.” — Fabian Hedin, Co-founder of Lovable Lovable ModelML Triple Whale PlayCo Canva Microsoft Base44 Legora Figma Pushing the frontier of cybersecurity and science
GPT‑5.6 is our strongest cybersecurity model yet, delivering frontier-level performance with fewer tokens. On ExploitBench2, which measures the full process from reaching vulnerable code to achieving arbitrary code execution, it scored 73.5% at a similar output-token budget, compared with 47.9% for GPT‑5.5. On ExploitGym3, where the benchmark requires an agent to turn real-world vulnerabilities into working exploits, GPT‑5.6’s best pass rate under a two-hour limit rose from GPT‑5.5’s 15.1% to 24.9%, nearly twice as high; under a six-hour limit, its pass rate reached 33.7%. On SEC-Bench Pro, which tests code generation for complex software proof-of-concepts (PoCs), it scored 71.2% with lower latency, compared with 45.8% for GPT‑5.5.1
GPT‑5.6 supports important defensive tasks including secure code review, patch development, threat modeling, and blue-team defense. Through OpenAI Daybreak’s Trusted Access for Cyber program, eligible individuals and organizations can gain stronger defensive capabilities with precise safeguards for verified work in authorized environments. This includes vulnerability triage and validation, malware analysis, detection engineering, and patch validation.
Individuals can verify their identity and apply for trusted access (Opens in a new window) , and organizations can apply for their teams. Individual members must enable Advanced Account Security using hardware-backed passkeys (Opens in a new window) by September 1 to retain access to our frontier models with strong cybersecurity performance; those who do not will return to default access. Users who do not yet have hardware-backed passkeys can receive an exclusive discount from our partner Yubico (Opens in a new window) . We are also taking additional measures to limit access by high-risk entities and from high-risk jurisdictions.
ExploitBench ExploitGym SEC-Bench Pro Capture the flag (CTF)
ExploitBench: Building progressively advanced V8 exploits; GPT‑5.6 shows a large improvement over GPT‑5.5. No latency chart is shown because latency estimates for this benchmark are unreliable.
GPT‑5.6 Sol also shows broad capability gains in scientific research. In life-sciences evaluations covering real-world biology, life-science research workflows, and chemistry, GPT‑5.6 shows Pareto improvements over GPT‑5.5.
GeneBench Pro LifeSciBench MedChemBench
GeneBench Pro: Long-horizon genomics and quantitative-biology analysis; GPT‑5.6 achieved better results with fewer tokens and in less time. Claude Fable 5 was not included because it does not answer advanced biology questions (Opens in a new window) and refused most questions in this evaluation.
GPT‑5.6 accelerates OpenAI
GPT‑5.6 is our strongest model yet for accelerating AI research. Inside OpenAI, researchers use it throughout the development loop: diagnosing failures, optimizing training systems, running experiments, and analyzing results. During internal GPT‑5.6 testing, we saw this acceleration trend alongside stronger adoption: daily output tokens per active researcher were more than twice the highest level observed for GPT‑5.5.
This way of working is quickly becoming the norm. Over the past six months, the share of research compute used for internal code reasoning has grown 100-fold, while internal agent-token usage has increased by about 22-fold. These usage metrics do not measure research progress by themselves, but they show that AI assistance is rapidly expanding in research and in other teams such as sales, marketing, user operations, and finance.
To measure this capability directly, we developed an internal evaluation suite based on real AI research tasks, covering debugging research systems, optimizing kernels and training recipes, running machine-learning experiments, and optimizing other models.
RSI Index Internal Research Debugging Evaluation KernelGen 1P NanoGPT
Overall RSI capability: Across a series of evaluations measuring progress in recursive self-improvement, we observed a 16.2-point gain for GPT‑5.6 Sol over GPT‑5.5, accelerating internal research across the board.
Safety and security that scale with model capability
As model capabilities continue to improve, we are also strengthening our safety systems so that advanced intelligence can remain broadly useful while the highest-risk uses receive more rigorous review. For GPT‑5.6, we built our most robust safety system yet, calibrated to each model’s capabilities and backed by substantial compute.
In biology and cybersecurity, GPT‑5.6 models exceed our earlier models, but neither domain has crossed the “Critical” risk threshold. In cybersecurity, our testing shows that GPT‑5.6 is better at finding and fixing vulnerabilities than at reliably carrying out autonomous, end-to-end attacks against hardened targets — giving defenders an opportunity to harden systems before vulnerabilities are exploited. In biology, our tests show that GPT‑5.6 can support legitimate scientific research, but it does not yet have the end-to-end capability required to create, design, or synthesize novel, highly dangerous threats.
Both domains are inherently dual-use. In cybersecurity, the same capability that can help an attacker exploit a vulnerability can help a defender discover and reproduce it and build a reliable fix. Overblocking therefore creates security risks of its own. It can prevent defenders from testing systems and deploying patches while malicious actors continue using other models (including increasingly capable open models) and existing hacking tools. Effective safeguards should account for the specific context and possible consequences of a request; they should protect legitimate defensive work while applying stricter controls when there is evidence of a serious risk of harm.
GPT‑5.6’s safeguards use a layered design to improve accuracy and redundancy and to adapt quickly to emerging attack methods. Model-level defenses work with real-time checks, continuous monitoring, and account-level interventions so that the overall system remains safe even if one layer does not work as intended. In many systems, a classifier flag alone determines what to block, often relying on less intelligent models that are difficult to update quickly to prevent harm. Our approach introduces a reasoning monitor that reviews conversations to determine whether potential harm is present. This design aims to keep defensive work moving while blocking serious abuse; the most sensitive capabilities are available only to verified users through the Trusted Access for Cyber program. Because some safeguards use test-time reasoning, we can update them quickly to close gaps without retraining a classifier from scratch.
As we continue strengthening the system against adaptive attacks, we have taken a more conservative approach. Compared with previous models, our cybersecurity safeguards for GPT‑5.6 Sol block roughly ten times as many potentially harmful activities. Because these measures may hinder legitimate users, ChatGPT and Codex provide an option to retry a prompt on a lower-capability model. We will continue reducing the impact on legitimate users while maintaining a high standard of system robustness. This reflects our iterative deployment strategy: start conservatively and keep optimizing based on lessons from real-world use.
Before the broad launch, we conducted an especially intensive safety evaluation. This included extensive red teaming, robustness and safeguard testing with external experts, and black-box automated red teaming equivalent to approximately 700,000 NVIDIA A100 Tensor Core GPU hours. This allowed us to systematically probe potential weaknesses, find jailbreaks, and strengthen the system before release.
Absolute safety does not exist, and our work to secure increasingly capable models never stops. New weaknesses continue to be found, and new jailbreak techniques for bypassing existing safeguards continue to emerge. Every new model generation brings new possibilities for attack and abuse. To address this reality, we have built layered defenses, continuous monitoring, rapid remediation processes, and active collaboration with the broader defense community. For GPT‑5.6, we combined our existing security (Opens in a new window) and biosecurity bug bounty programs with a new rapid-remediation process and our most rigorous monitoring to date. Findings from researchers, system monitoring, and real-world abuse cases will continue to become new evaluation benchmarks and stronger safeguards.
For more information about our safeguards, see the updated GPT‑5.6 System Card (Opens in a new window) .
Availability and pricing
GPT‑5.6 is available in three model tiers: our flagship Sol; Terra, with performance comparable to GPT‑5.5 at lower cost; and faster, highly cost-effective Luna. The number identifies the generation, while Sol, Terra, and Luna are persistent capability tiers that can evolve independently at their own pace.
Starting today, GPT‑5.6 is available in ChatGPT, Codex, and the OpenAI API. The series is beginning its global rollout and will become available to all users over the next 24 hours.
Chat: Plus, Pro, Business, and Enterprise users can access GPT‑5.6 Sol with medium and higher reasoning-effort settings. Pro and Enterprise users can also select GPT‑5.6 Sol Pro for the highest-quality results on complex tasks. ChatGPT Work and Codex: Free and Go users can access GPT‑5.6 Terra. Plus, Pro, Business, and Enterprise users can choose among GPT‑5.6 Sol, Terra, and Luna and set reasoning effort for each. Everyone with GPT‑5.6 access in ChatGPT Work and Codex can use the max option, which can be enabled in settings. In ChatGPT Work, ultra is available to Pro and Enterprise users. In Codex, it is available to Plus and higher subscription tiers. API: Developers can access Sol, Terra, and Luna through the OpenAI API. In the Responses API, programmable tool calling lets GPT‑5.6 write and run programs in memory to coordinate tools and process intermediate results in a way that meets the “zero data retention” (ZDR) standard. The multiagent feature, currently in testing, lets GPT‑5.6 run subagents in parallel and combine their work in a single request.
GPT‑5.6 comes in three model sizes, priced per million (1M) tokens: Sol costs $5 for input and $30 for output; Terra costs $2.50 for input and $15 for output; Luna costs $1 for input and $6 for output. GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints (Opens in a new window) and a cache lifetime of at least 30 minutes. For GPT‑5.6 and later models, cache writes cost 1.25 times the model’s uncached input rate, while cache reads continue to receive a 90% discount from the cached-input rate.
7, 8
Professional capabilities Evaluation GPT‑5.6 Sol GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Fable 5 Claude Opus 4.8 Gemini 3.1 Pro Preview Gemini 3.5 Flash Agents' Last Exam 52.7% 50.4% 50.3% 46.9% 40.5% 45.2% 32.1% — GDPval-AA v2 1,747.8 Elo 1,593 Elo 1,591.8 Elo 1,493.7 Elo 1,759.6 Elo 1,600.1 Elo 962.3 Elo 1,348.8 Elo Management consulting tasks (internal) 43.2% 37.2% 35.4% 31.3% 35.5% 31.6% 13.2% — Big Finance Bench 53% 51% 36% 49% — 44% — — Artificial Analysis Intelligence Index v4.1 58.9 Index score 55 Index score 51.2 Index score 54.8 Index score 59.9 Index score 55.7 Index score 46.5 Index score 50.2 Index score Coding Evaluation GPT‑5.6 Sol GPT‑5.6 Sol Ultra GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Mythos 5 Claude Mythos Preview Claude Fable 5 Claude Opus 4.8 Gemini 3.1 Pro Preview Artificial Analysis Coding Agent Index v1.1 80 Index score — 77.4 Index score 74.6 Index score 76.4 Index score — — 77.2 Index score 72.5 Index score 42.7 Index score SWE-Bench Pro 64.6% — 63.4% 62.7% 59.4% 80.3% 77.8% 80% 69.2% 54.2% DeepSWE v1.1 72.7% — 69.6% 67.2% 67% — — 69.7% 59% 11.8% Terminal-Bench 2.1 88.8% 91.9% 87.4% 84.7% 85.6% 88% — 83.1% 78.9% 70.7% Safety Evaluation GPT‑5.6 Sol GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 GPT‑5.4 Claude Opus 4.8 Claude Mythos 5 Claude Mythos Preview Healthbench Professional 60.5% 57.7% 55.7% 51.8% 48.1% 52.6% 66% 64.7% Computer use Evaluation GPT‑5.6 Sol GPT‑5.6 Sol Ultra GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Mythos 5 Claude Mythos Preview Claude Opus 4.8 Gemini 3.1 Pro Preview OSWorld 2.0 62.6% — 50.2% 45.6% 47.5% — — 54.8% — BrowseComp 90.4% 92.2% 87.5% 83.3% 84.4% 88% 87.9% 84.3% 85.9% BenchCAD 70.6% — 62.3% 63.1% 44.4% 38.4% 35.5% 27.3% — BenchCAD (Python tool) 83.4% — 78.2% 73.9% 55.8% 65% 61% 51.8% — Cybersecurity Evaluation GPT‑5.6 Sol GPT‑5.6 Sol Ultra GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Mythos 5 Claude Mythos Preview Claude Opus 4.8 Capture the flag challenge 96.7% — 91.8% 85.2% 88.1% — — — SEC-Bench Pro 71.2% 74.3% 57.7% 48.9% 45.8% — — — ExploitBench 73.5% — 52.9% 33.2% 47.9% 78% 74.2% 40% ExploitGym 33.7% — 23.2% 12.4% 15.1% — — — Self-improvement Evaluation GPT‑5.6 Sol GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Internal Research Debugging Evaluation 68.3% 67.8% 50.8% 50% KernelGen 1P 61.1% 49.2% 22.4% 29.3% NanoGPT 9.69% 14.5% 1.66% 2.65% PostTrainBench Lite 50.3% 51.5% 29.6% 38.8% RSI Index 57.9% 56.3% 41.9% 41.7% Multimodal Evaluation GPT‑5.6 Sol GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Fable 5 Claude Opus 4.8 Gemini 3.1 Pro Preview MMMU Pro (without tools) 83% 80.7% 78.4% 81.2% — — 80.5% MMMU Pro (with tools) 84.6% 82% 79.5% 83.2% — — — gdp.pdf 30.7% 24.7% 22.7% 26% 29.8% 22.5% 16.7% Academic Evaluation GPT‑5.6 Sol GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Mythos 5 Claude Mythos Preview Claude Fable 5 Claude Opus 4.8 Gemini 3.1 Pro Preview GPQA Diamond 94.6% 92.9% 92.3% 93.6% 94.1% 94.6% 92.6% 92% 94.3% FrontierMath Tier 1-3 (v2) 89% 84.9% 78.6% 85.3% — — 87% 80% 59.6% FrontierMath Tier 4 (v2) 83% 68.3% 58.5% 72.5% — — 87.8% 56.1% — Tool use Evaluation GPT‑5.6 Sol GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Mythos 5 Claude Mythos Preview Claude Fable 5 Claude Opus 4.8 Gemini 3.1 Pro Preview Gemini 3.5 Flash AutomationBench 18.1% 15.2% 14.9% 12.9% — — 17.4% 15.5% — 14.5% Toolathlon 58% 53.1% 53.4% 55.6% 61.7% 61.1% 61.7% 59.9% 48.8% — Long context Evaluation GPT‑5.6 Sol GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Mythos 5 Claude Mythos Preview Claude Opus 4.8 OpenAI MRCR v2 8-needle 256K-512K 91.5% 89.6% 41.3% 81.5% — — — OpenAI MRCR v2 8-needle 512K-1M 73.8% 72.5% 41.3% 74% — — — GraphWalks BFS 256k f1 90.7% 76.9% 81.3% 73.7% 91.1% 85.7% 85.9% GraphWalks BFS 1mil f1 77.1% 71.2% 51.2% 45.4% 79.4% 74.3% 68.1% Abstract reasoning Evaluation GPT‑5.6 Sol GPT‑5.6 Terra GPT‑5.6 Luna GPT‑5.5 Claude Opus 4.8 Gemini 3.1 Pro Preview ARC-AGI-3⁷ 7.78% 0.8% 0.18% 0.43% 1.5% 0.42%
Translation feedback
Was this page easy to read and understand?
Excellent Good Poor 2026 Author OpenAI Footnotes 1
The cybersecurity evaluation was conducted with relaxed safeguard restrictions. Users can join OpenAI Daybreak’s Trusted Access for Cyber program to gain access to additional cybersecurity defense capabilities.
2
All models were evaluated using the ExploitBench API test harness with 5 seeds and reasoning continuity enabled.
3
We ran ExploitGym on an alpha API whose response-generation speed is faster than the public API, then rescaled it to match the public API. When latency is rescaled to the expected public API speed, some estimated latencies exceed the two- and six-hour limits, even though those limits were followed in the evaluation runs. For time-sensitive work that needs faster speed, we offer priority processing in the API and fast mode in Codex.
4
We estimate latency and API cost by observing model performance in production and running offline simulations. These estimates account for tool-call details, sampled tokens, and input tokens. Actual results can vary substantially and depend on many factors not covered by our simulations. We simulate latency at high-speed API rates and cost at standard API pricing.
5
Models that did not report output tokens, latency, or cost are shown with horizontal dashed lines.
6
For multi-agent systems, latency is calculated from the root agent, while output tokens and total API cost include all tokens. Ultra mode runs with a four-agent configuration.
7
We used the official scoring method described in the HealthBench Professional paper. The resulting score is not comparable with the result published in Anthropic’s system card.
8
The ARC-AGI-3 test for Opus 4.8 was run at high rather than max reasoning effort because this is currently the model’s only public ARC-AGI-3 result.
Read more View all Ultrafast mode preview: GPT-5.6 Sol reaches up to 14x faster speeds
Products August 13, 2026
Test ads in ChatGPT
Company August 11, 2026
Daybreak models are now available on AWS
Products August 11, 2026
Research Research index Research overview Economic research Latest advances GPT-5.6 GPT-5.5 GPT-5.4 Safety Safety measures Deployment safety (Opens in a new window) Security and privacy Trust and transparency Products ChatGPT (Opens in a new window) ChatGPT Business (Opens in a new window) ChatGPT Enterprise (Opens in a new window) ChatGPT for Education (Opens in a new window) Codex Release notes API Platform Overview API login (Opens in a new window) Documentation (Opens in a new window) Business Overview Solutions Resources Customer stories Partner network Contact sales Developers Apps SDK (Opens in a new window) Open models Documentation (Opens in a new window) Resources (Opens in a new window) Developer forum (Opens in a new window) Company About us Our charter Careers News Support Help center (Opens in a new window) More Customer stories Academy Supply Co. Live Podcast RSS Terms and policies Terms of use Privacy policy Other policies (Opens in a new window) (Opens in a new window) (Opens in a new window) (Opens in a new window) (Opens in a new window) (Opens in a new window) (Opens in a new window) OpenAI © 2015–2026 Your privacy choices Chinese China
GPT-5.6 Sol