Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Sol · Media / benchmark · Independent measurement

Artificial Analysis: Sol's intelligence, coding-agent result, and cost per task

Artificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Model version
GPT-5.6 Sol max; also compares Terra, Luna, and GPT-5.5
Provider / client
Artificial Analysis; Sol runs in OpenAI Codex for Coding Agent
Reasoning tier
max; other GPT-5.6 tiers are also shown
Tools
Codex coding-agent harness; other models use their associated harnesses
Task set
Intelligence Index v4.1; Coding Agent (DeepSWE, Terminal-Bench v2, SWE-Atlas-QnA); AA-Briefcase
Sample / repeats
Complete sample size and repeat count not disclosed
Publication / collection date
2026-07-09 / 2026-08-17
Traceable results
Intelligence 59; Coding Agent 80; about $1.04/task; about 15,000 output tokens/task

Key data and applicable tasks

Summary

Third-party benchmarks record GPT-5.6 Sol's performance on the Intelligence Index and Coding Agent Index, comparing its scores, cost per task, and latency with models including Claude Fable 5.

Original article

The following is the visible body text extracted during this visit. It includes page navigation, machine translation, advertising, comments, and other page elements; verify against the original link before citing it.


Artificial Analysis Models Coding Agents Speech, Image, Video Inference Leaderboards About AI Trends Arenas Premium Log in K All articles

July 9, 2026

GPT-5.6 benchmarks across Intelligence, Speed and Cost See model page

GPT-5.6 Sol comes close second to Claude Fable 5 in the Artificial Analysis Intelligence Index at one third of the cost, and leads the Artificial Analysis Coding Agent Index in OpenAI’s Codex harness

We supported OpenAI with pre-release evaluation of GPT-5.6 Sol, Terra, and Luna. GPT-5.6 Sol (max) scores 1 point below Claude Fable 5 (max) in the Artificial Analysis Intelligence Index at 59 points, at approximately one third of the cost. GPT-5.6 Terra (max) and Luna (max) score 55 and 51 respectively in the Intelligence Index, at ~50% and ~80% lower Cost per Task than Sol.

GPT-5.6 Sol (max) leads the Artificial Analysis Coding Agent Index at 80 points.

Key takeaways:

➤ One third of the cost of Claude Fable 5: On max reasoning effort, GPT-5.6 Sol costs $1.04 per task in the Artificial Analysis Intelligence Index - offering a similar level of intelligence to Claude Fable 5 at approximately one third of the cost. Reasoning levels across GPT-5.6 Sol and Luna offer a range of options at the Pareto frontier of Intelligence vs Cost per Task. For example, GPT-5.6 Luna (max) matches or exceeds the intelligence of GLM-5.2 (max) and Gemini 3.5 Flash at a lower cost. GPT-5.6 Terra (max) and Luna (max) cost $0.55 and $0.21 per Intelligence Index task, ~50% and ~80% less than Sol. Across reasoning efforts, each new GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excluding non-reasoning). Notably, Luna and Sol are always on the Pareto frontier ahead of Terra. This means that for any Terra effort level, there is a Luna or Sol effort level that is more intelligent at no extra cost, or equally intelligent at lower cost.

➤ Leading in all Coding Agent evaluations: The new Artificial Analysis Coding Agent Index pairs models with agentic harnesses and features three frontier coding evaluations - DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. GPT-5.6 Sol (max) in Codex scores 80 in the Index, leading in all three evaluations (tying Grok 4.5 in Grok Build for SWE-Atlas-QnA). In addition to scoring higher, its per task cost is ~40% and ~10% cheaper than Claude Fable 5 (max) and Opus 4.8 (max) respectively in Claude Code. GPT-5.6 Terra (max) and Luna (max) score 77 and 75 in the Coding Agent Index respectively, with ~60% and ~80% per-task cost reductions compared to Sol.

➤ Highest Presentation Elo in AA-Briefcase: GPT-5.6 Sol (max) ranks second only to Claude Fable 5 (max) in AA-Briefcase, and has the highest Presentation Elo of any model. AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. GPT-5.6 Sol (max) has the highest recorded Presentation Elo - its outputs across various file types, including PowerPoint and Excel, are the most visually attractive of any model. Fable 5 (max) still leads AA-Briefcase, largely due to its Rubric Score of 56% vs 42% for GPT-5.6 Sol (max). Fable 5 (max) also scores 1764 in Analytical Quality Elo vs GPT-5.6 Sol (max) at 1592.

➤ First OpenAI models with cache-write pricing: GPT-5.6 introduces cache-write pricing for the first time at OpenAI. Sol, Terra, and Luna are priced at $5/$30, $2.5/$15, and $1/$6 respectively per million input/output tokens. OpenAI has retained its previous discount of 90% for cache reads, but joins Anthropic in introducing a cost premium for cache writes, at 1.25x the price of input tokens. Cache writes occur when input tokens are committed to memory. Charging for a cache write more accurately reflects the model’s cost to serve, as cached tokens occupy memory whether or not they are reused. Also in line with Anthropic’s models, GPT-5.6 introduces a max reasoning effort level.

➤ Low token use: GPT-5.6 Sol (max) uses fewer output tokens than most models of comparable intelligence, and defines a new Pareto frontier of Intelligence vs Output Tokens per Task. GPT-5.6 Sol (max) offers a slight improvement in token efficiency with 15k tokens per Intelligence Index task, vs GPT-5.5 at 16k. Notably, it uses fewer tokens and is more intelligent than Claude Opus 4.8 (max), GLM-5.2 (max), and Gemini 3.5 Flash (high).

GPT-5.6 Sol (max) offers a similar level of intelligence to Claude Fable 5 at approximately one third of the cost. The model family defines a new Pareto frontier of Intelligence vs Cost per Task.

Across reasoning efforts, each GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excludes non-reasoning). Notably, Luna and Sol are always on the Pareto frontier ahead of Terra.

GPT-5.6 Sol (max) in Codex leads every evaluation in the Artificial Analysis Coding Agent Index (tying Grok 4.5 in Grok Build for SWE-Atlas-QnA). It has lower Cost per Task than Claude Fable 5 (max) and Claude Opus 4.8 (max).

GPT-5.6 Sol (max) ranks second only to Claude Fable 5 (max) in AA-Briefcase, and has the highest Presentation Elo of any model.

GPT-5.6 Sol defines a new Pareto frontier of Intelligence vs Output Tokens per Task in the Artificial Analysis Intelligence Index. Terra and Luna are not on the Pareto frontier.

GPT-5.6 Sol (max) scores similarly to Claude Fable 5 (max) in GDPval-AA v2, reflecting a similar ability to complete economically valuable tasks.

GPT-5.6 Sol (max) offers a minor improvement over GPT-5.5 in the AA-Omniscience Index, with a small uplift in accuracy coupled with an increase in hallucination rate.

Breakdown of the individual evaluations in the Artificial Analysis Intelligence Index v4.1.

Compare GPT-5.6 Sol, Terra, and Luna with other leading models at: https://artificialanalysis.ai

Newsletter

Get notified about new articles Email address Subscribe

We'll email you when we publish something new.

Read the latest Announcing Optima: create a custom benchmark for your use case

Optima is a new platform for benchmarking models on your own workloads. Build a benchmark from your own files, agent traces or coding environment, run it across leading models in a single click, and compare quality alongside cost per task and time per task.

August 13, 2026

Gemini 3.7 Flash: On the Intelligence vs. Time per Task Pareto frontier

Google has released Gemini 3.7 Flash, improving 4 points over Gemini 3.6 Flash and reaching the Intelligence vs. Time per Task Pareto frontier

August 13, 2026

Upstage Solar Pro 4: Benchmarks and analysis

Upstage has released Solar Pro 4

August 12, 2026

Get notified about new articles

Email address Subscribe

Artificial Analysis

Explore

LLM Leaderboard Image Arena Video Arena AI Agents Evaluations

Company

Methodology Services Contact Articles FAQ X LinkedIn YouTube Rednote Discord

© 2026 Artificial Analysis

Terms of Use Privacy Policy

What this supports

  • Supports comparing Sol max's intelligence, cost per task, output tokens, and coding-agent result within this evaluation's methodology.
  • Supports distinguishing end-to-end harness results such as Codex from model weights alone.

What this does not support

  • Does not support extrapolating 59, 80, or cost per task to other clients, live prices, or all tasks.
  • Artificial Analysis supported pre-release OpenAI evaluation, so its design and sample limits remain material.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis · Original publication date 2026-07-09 · Site edit date 2026-09-20

Open original source

GPT-5.6 Sol

Compare GPT-5.6 Sol in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Sol: Specs, Access, Changes, and the Risks That Still Matter

OpenAI's current GPT-5.6 Sol model page lists a 1.05M context window, 128K max output, reasoning controls, and a time-sensitive API price card. Here is what those facts mean for API, Codex, and browser users.

Related reviews

Visual Studio Magazine: Sol's token efficiency and reasoning-slider limitsThe article describes one Sol model behind quick and deeper Plus/Pro responses and cites 80 on Coding Agent Index, 64.6% on SWE-Bench Pro, and about 15,000 output tokens per Intelligence task; Sol did not lead every evaluation.METR: Sol's time horizon changes with cheating treatmentIn Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.OpenAI release note: Sol's official results on long-horizon, coding, and knowledge workOpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.Every: Sol excels as a collaborative knowledge-work partner, not as judgmentEvery describes Sol as fast and steerable across 24 drafts, email, meetings, and retrieval, but it scored 56/100 versus Fable's 90/100 on Senior Engineer and ranked last of six in writing; collaboration is not autonomous judgment.Route ChatGPT tasks through Sol’s reasoning settingsRun the same task at faster and deeper reasoning settings: prefer speed for short questions, then increase reasoning for planning, research, writing, coding, and decisions; OpenAI’s 68% figure is an internal relative change, not public accuracy.Design a verifiable multi-agent workflow with the Responses APISeparate judgment from deterministic processing, then combine programmatic tool calls, parallel subagents, and prompt-cache boundaries into a long-running workflow whose cost, latency, citations, and failures can be reviewed.Configure Codex for a million-token context and auto-compactionThe source shows config.toml and one-session CLI examples for the model ID, a 1,000,000-token context budget, and a 900,000-token compaction threshold; confirm client support and keep a rollback configuration before editing.Deliver code with prediction, planning, review, and verificationSplit long-running coding into prediction, planning, implementation, adversarial review, and independent verification, checking the plan, tests, and stop conditions item by item; this is a commenter’s personal workflow, not Codex’s default configuration.