Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaGPT-5.6 Sol

GPT-5.6 benchmarks across Intelligence, Speed and Cost

Original source

Artificial Analysis

AuthorArtificial Analysis

Source date2026-07-09

Tabbit curation2026-08-19

Read original

Summary

Third-party benchmarks record GPT-5.6 Sol's performance on the Intelligence Index and Coding Agent Index, comparing its scores, cost per task, and latency with models including Claude Fable 5.

Original article

The following is the visible body text extracted during this visit. It includes page navigation, machine translation, advertising, comments, and other page elements; verify against the original link before citing it.


Artificial Analysis Models Coding Agents Speech, Image, Video Inference Leaderboards About AI Trends Arenas Premium Log in K All articles

July 9, 2026

GPT-5.6 benchmarks across Intelligence, Speed and Cost See model page

GPT-5.6 Sol comes close second to Claude Fable 5 in the Artificial Analysis Intelligence Index at one third of the cost, and leads the Artificial Analysis Coding Agent Index in OpenAI’s Codex harness

We supported OpenAI with pre-release evaluation of GPT-5.6 Sol, Terra, and Luna. GPT-5.6 Sol (max) scores 1 point below Claude Fable 5 (max) in the Artificial Analysis Intelligence Index at 59 points, at approximately one third of the cost. GPT-5.6 Terra (max) and Luna (max) score 55 and 51 respectively in the Intelligence Index, at ~50% and ~80% lower Cost per Task than Sol.

GPT-5.6 Sol (max) leads the Artificial Analysis Coding Agent Index at 80 points.

Key takeaways:

➤ One third of the cost of Claude Fable 5: On max reasoning effort, GPT-5.6 Sol costs $1.04 per task in the Artificial Analysis Intelligence Index - offering a similar level of intelligence to Claude Fable 5 at approximately one third of the cost. Reasoning levels across GPT-5.6 Sol and Luna offer a range of options at the Pareto frontier of Intelligence vs Cost per Task. For example, GPT-5.6 Luna (max) matches or exceeds the intelligence of GLM-5.2 (max) and Gemini 3.5 Flash at a lower cost. GPT-5.6 Terra (max) and Luna (max) cost $0.55 and $0.21 per Intelligence Index task, ~50% and ~80% less than Sol. Across reasoning efforts, each new GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excluding non-reasoning). Notably, Luna and Sol are always on the Pareto frontier ahead of Terra. This means that for any Terra effort level, there is a Luna or Sol effort level that is more intelligent at no extra cost, or equally intelligent at lower cost.

➤ Leading in all Coding Agent evaluations: The new Artificial Analysis Coding Agent Index pairs models with agentic harnesses and features three frontier coding evaluations - DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. GPT-5.6 Sol (max) in Codex scores 80 in the Index, leading in all three evaluations (tying Grok 4.5 in Grok Build for SWE-Atlas-QnA). In addition to scoring higher, its per task cost is ~40% and ~10% cheaper than Claude Fable 5 (max) and Opus 4.8 (max) respectively in Claude Code. GPT-5.6 Terra (max) and Luna (max) score 77 and 75 in the Coding Agent Index respectively, with ~60% and ~80% per-task cost reductions compared to Sol.

➤ Highest Presentation Elo in AA-Briefcase: GPT-5.6 Sol (max) ranks second only to Claude Fable 5 (max) in AA-Briefcase, and has the highest Presentation Elo of any model. AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. GPT-5.6 Sol (max) has the highest recorded Presentation Elo - its outputs across various file types, including PowerPoint and Excel, are the most visually attractive of any model. Fable 5 (max) still leads AA-Briefcase, largely due to its Rubric Score of 56% vs 42% for GPT-5.6 Sol (max). Fable 5 (max) also scores 1764 in Analytical Quality Elo vs GPT-5.6 Sol (max) at 1592.

➤ First OpenAI models with cache-write pricing: GPT-5.6 introduces cache-write pricing for the first time at OpenAI. Sol, Terra, and Luna are priced at $5/$30, $2.5/$15, and $1/$6 respectively per million input/output tokens. OpenAI has retained its previous discount of 90% for cache reads, but joins Anthropic in introducing a cost premium for cache writes, at 1.25x the price of input tokens. Cache writes occur when input tokens are committed to memory. Charging for a cache write more accurately reflects the model’s cost to serve, as cached tokens occupy memory whether or not they are reused. Also in line with Anthropic’s models, GPT-5.6 introduces a max reasoning effort level.

➤ Low token use: GPT-5.6 Sol (max) uses fewer output tokens than most models of comparable intelligence, and defines a new Pareto frontier of Intelligence vs Output Tokens per Task. GPT-5.6 Sol (max) offers a slight improvement in token efficiency with 15k tokens per Intelligence Index task, vs GPT-5.5 at 16k. Notably, it uses fewer tokens and is more intelligent than Claude Opus 4.8 (max), GLM-5.2 (max), and Gemini 3.5 Flash (high).

GPT-5.6 Sol (max) offers a similar level of intelligence to Claude Fable 5 at approximately one third of the cost. The model family defines a new Pareto frontier of Intelligence vs Cost per Task.

Across reasoning efforts, each GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excludes non-reasoning). Notably, Luna and Sol are always on the Pareto frontier ahead of Terra.

GPT-5.6 Sol (max) in Codex leads every evaluation in the Artificial Analysis Coding Agent Index (tying Grok 4.5 in Grok Build for SWE-Atlas-QnA). It has lower Cost per Task than Claude Fable 5 (max) and Claude Opus 4.8 (max).

GPT-5.6 Sol (max) ranks second only to Claude Fable 5 (max) in AA-Briefcase, and has the highest Presentation Elo of any model.

GPT-5.6 Sol defines a new Pareto frontier of Intelligence vs Output Tokens per Task in the Artificial Analysis Intelligence Index. Terra and Luna are not on the Pareto frontier.

GPT-5.6 Sol (max) scores similarly to Claude Fable 5 (max) in GDPval-AA v2, reflecting a similar ability to complete economically valuable tasks.

GPT-5.6 Sol (max) offers a minor improvement over GPT-5.5 in the AA-Omniscience Index, with a small uplift in accuracy coupled with an increase in hallucination rate.

Breakdown of the individual evaluations in the Artificial Analysis Intelligence Index v4.1.

Compare GPT-5.6 Sol, Terra, and Luna with other leading models at: https://artificialanalysis.ai

Newsletter

Get notified about new articles Email address Subscribe

We'll email you when we publish something new.

Read the latest Announcing Optima: create a custom benchmark for your use case

Optima is a new platform for benchmarking models on your own workloads. Build a benchmark from your own files, agent traces or coding environment, run it across leading models in a single click, and compare quality alongside cost per task and time per task.

August 13, 2026

Gemini 3.7 Flash: On the Intelligence vs. Time per Task Pareto frontier

Google has released Gemini 3.7 Flash, improving 4 points over Gemini 3.6 Flash and reaching the Intelligence vs. Time per Task Pareto frontier

August 13, 2026

Upstage Solar Pro 4: Benchmarks and analysis

Upstage has released Solar Pro 4

August 12, 2026

Get notified about new articles

Email address Subscribe

Artificial Analysis

Explore

LLM Leaderboard Image Arena Video Arena AI Agents Evaluations

Company

Methodology Services Contact Articles FAQ X LinkedIn YouTube Rednote Discord

© 2026 Artificial Analysis

Terms of Use Privacy Policy

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-5.6 Sol

Use and compare models in Tabbit

GPT-5.6 Sol

Related reviews

OfficialOpenAI2026-07-09

GPT-5.6: Frontier Intelligence That Scales Flexibly to Ambitious Goals

OfficialOpenAI Deployment Safety Hub2026-07-09

OpenAI GPT‑5.6 System Card: Safety, Prompt Injection, and Agent Boundaries

MediaCodeRabbit2026-07-09

OpenAI GPT-5.6 Sol and Terra: Benchmark

MediaVisual Studio Magazine2026-08-06

GPT-5.6 Sol Ascends for Token Efficiency; How Does It Stack Up Against Other Models?

GPT-5.6 Sol

Related prompts

OfficialOpenAI2026-08-13

The builder’s guide to GPT‑5.6

OfficialOpenAI2026-08-06

GPT‑5.6 Sol: ChatGPT Reasoning Slider and Task Routing Configuration

OfficialOpenAI2026-08-13

GPT-5.6 Sol Ultrafast: Real-time Workflow Configuration and Integration Boundaries

CommunityThe Prompt Index

GPT-5.6 (Sol) & Claude Fable 5 Prompting Guide (2026)