Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Sol · Media / benchmark · Editorial analysis

Visual Studio Magazine: Sol's token efficiency and reasoning-slider limits

The article describes one Sol model behind quick and deeper Plus/Pro responses and cites 80 on Coding Agent Index, 64.6% on SWE-Bench Pro, and about 15,000 output tokens per Intelligence task; Sol did not lead every evaluation.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Model version
GPT-5.6 Sol; August ChatGPT update separated from the July evaluation
Provider / client
OpenAI ChatGPT Plus/Pro web account; author used a company account
Reasoning tier
Sol Light and a deeper slider; cited benchmark tiers are separate
Tools
ChatGPT web conversations; cited Coding Agent results use their own harnesses
Task set
Financial/medical/legal factual-answer evaluation; Artificial Analysis; SWE-Bench Pro
Sample / repeats
Internal samples not disclosed; cited benchmark repeat design not fully described
Publication / collection date
2026-08-06 / 2026-08-17
Traceable results
Coding Agent 80; SWE-Bench Pro 64.6%; about 15,000 tokens/task; OpenAI reports about 68% fewer errors

Key data and applicable tasks

Summary

The article draws on real-world use of ChatGPT Plus/Pro and Artificial Analysis data to discuss Sol’s token efficiency, quick and deep response modes, and the fact that it does not lead every evaluation.

Original article

The following is the visible body text extracted during this visit. It includes page navigation, machine translation, advertising, comments, and other page elements; verify against the original link before citing it.


Skip to main content Add as a preferred source on Google HOME NEWS TIPS & HOW-TO NEWSLETTERS WHITE PAPERS WEBCASTS ADVERTISE ABOUT US TRAINING MORE AZUREVISUAL STUDIOVISUAL STUDIO CODEBLAZOR/ASP.NET.NETC#/VB/TYPESCRIPT.NET MAUI/MOBILEAI/MACHINE LEARNING

NEWS

GPT-5.6 Sol Ascends for Token Efficiency; How Does It Stack Up Against Other Models? By David Ramel08/06/2026 Key Takeaways GPT-5.6 Sol now powers quick and deeper responses for ChatGPT Plus and Pro. OpenAI says the new ChatGPT models replace GPT-5.5 Instant. Benchmark results show strong efficiency, but Sol does not lead every evaluation.

When I fired up my web-based, company-paid ChatGPT account recently, I noticed ChatGPT-5.6 Sol Light was my default model. I don't usually pay too much attention to what model I'm using on various systems throughout the day, but spotting Sol reminded me of some recent chatter about it being picked by Microsoft for internal usage because of its efficiency in using tokens.

And token consumption, of course, is in the news a lot lately, especially in the wake of the June 1 switch to usage-based billing for GitHub Copilot. Developers have been reporting rapid depletion of their AI Credit allowances, and Microsoft has reportedly told employees that maximizing AI use is no longer the company's goal. If you will remember, it wasn't that long ago that corporate America was advocating "tokenmaxxing" as a way to measure adoption and engagement with AI tools. Now, the emphasis is shifting toward how much useful work the system produces for the time, tokens and money consumed.

Poking around, I found that OpenAI has put GPT-5.6 Sol behind both quick answers and deeper reasoning in ChatGPT for Plus and Pro subscribers, replacing the split between an everyday Instant model and a separate higher-reasoning experience.

An Aug. 6 ChatGPT update introduces a slider that lets paid users control how much thought the model applies to a response. OpenAI said "the same model now powers both Instant responses and deeper reasoning," with lower settings intended for everyday questions and higher settings available for planning, research, writing, coding and other more involved work.

An accompanying GPT-5.6 system card states that the updated Sol model for Plus and Pro and the updated Luna model for lower-cost plans "will replace GPT-5.5 Instant." GPT-5.6 Luna will become the default for Free and Go users this week, followed by unlimited text chats and a Think button for more difficult questions.

From Adoption Metrics to Token Budgets

The timing places OpenAI's unified Sol experience against a wider change in how companies discuss AI consumption. Two days before the ChatGPT announcement, 404 Media reported on internal Microsoft guidance that introduced limits on how much engineers could spend on workplace AI tools. The report's headline quoted Microsoft's message as "Tokenmaxxing is not what we are optimizing for," while its publicly available introduction said maximizing internal AI use was no longer the company's goal.

That guidance marks a different emphasis from the adoption measurements that vendors have encouraged organizations to follow during AI rollouts. Microsoft's official Microsoft 365 Copilot usage report, for example, is designed to summarize how users adopt, retain and engage with Copilot. Its measurements include enabled users, active users, active-user rates, total prompts, average prompts per active user and active days.

Those metrics remain available and relevant to deployment decisions. What has changed is that activity can now be evaluated alongside a more explicit consumption ledger. Under GitHub's usage-based Copilot billing system, input tokens, generated output and cached context are priced according to the selected model and converted into GitHub AI Credits. One credit equals 1 cent.

The documents cover different products and do not by themselves prove a Microsoft-wide reversal in AI policy. They do show the two measurements organizations now have to reconcile: whether employees are actively adopting AI tools and how much model consumption that activity generates.

Visual Studio Magazine documented that tension around GitHub's June 1 billing transition, when developers objected to replacing premium requests with token-based AI Credits. Subsequent coverage described rapid credit consumption after the system went live, including agentic workflows that depleted monthly allowances much faster than users expected.

GitHub and Microsoft have since introduced additional usage displays, spending limits and token-reduction work. A July review of those changes found a growing collection of meters, caps and token-saving tools across Copilot, Visual Studio and Visual Studio Code.

The new ChatGPT experience is not the same billing arrangement. OpenAI's announcement does not say Plus and Pro subscribers will receive per-token bills for ordinary ChatGPT conversations. The connection is instead at the model-design and resource-allocation level: the amount of reasoning applied to a response can now vary within the same Sol model, while the lower-cost Luna model handles default access for Free and Go users.

What the Sol Slider Changes

OpenAI describes the update as an effort to make quick and more deliberate ChatGPT responses feel like different effort levels from the same system rather than exchanges with separate models that have different tones and behaviors.

The company said the updated Sol produces more focused answers, uses tighter formatting and avoids unnecessary detail. It also reported an internal evaluation covering financial, medical and legal prompts that required factual detail. Responses containing at least one factual error were about 68% less common with the updated Sol than with GPT-5.5 Instant, according to OpenAI.

Here is a comparison of two responses to the query, "Can I bike from the Mission to Ocean Beach after work today without getting soaked?", with OpenAI preferring the Sol answer.

[Click on image for larger view.] 'GPT-5.6 Sol is the stronger answer' (source: OpenAI).

"GPT-5.6 Sol is the stronger answer because it answers the real question first, identifies wind rather than rain as the main issue, and keeps only the details the rider needs," OpenAI said. "After the 5:30 follow-up, it updates the recommendation without repeating the full forecast."

That result is directly relevant to the August ChatGPT model, but it remains a vendor-run internal evaluation. OpenAI did not provide the full prompt set, individual outputs or an external reproduction with the announcement.

How GPT-5.6 Sol Stacks Up

When I started wondering about the performance of my new default model, one thing that came to mind was AI going rogue. That has been in the news lately as both Claude and some OpenAI models went off the rails.

It so happens OpenAI addressed that issue with a measurement of jailbreaks. "We evaluate model robustness to jailbreaks: adversarial prompts designed to circumvent model refusal training and elicit harmful assistance," the company said in the August updates PDF. "GPT-5.6 Sol and GPT-5.6 Luna perform comparably to recent predecessors," OpenAI said."

[Click on image for larger view.] Jailbreaking (source: OpenAI).

Broader comparisons are available from the original July GPT-5.6 release, which evaluated Sol, Terra and Luna against GPT-5.5 and models from Anthropic and Google. These results cover the previously released GPT-5.6 versions, not the ChatGPT-tuned August Sol update.

Model Coding Agent Index SWE-Bench Pro GPT-5.6 Sol 80 64.6% GPT-5.6 Terra 77.4 63.4% GPT-5.6 Luna 74.6 62.7% GPT-5.5 76.4 59.4% Claude Fable 5 77.2 80% Claude Opus 4.8 72.5 69.2% Gemini 3.1 Pro Preview 42.7 54.2%

Sol led that group on the Artificial Analysis Coding Agent Index, which combines several agentic coding evaluations. It also finished ahead of GPT-5.5, Terra, Luna, Claude Fable 5, Claude Opus 4.8 and Gemini 3.1 Pro Preview on that composite measurement.

The ranking changed on SWE-Bench Pro. Claude Fable 5 scored 80%, Claude Opus 4.8 scored 69.2% and Sol scored 64.6%. Sol still exceeded GPT-5.5, Terra, Luna and Gemini 3.1 Pro Preview, but it did not lead the evaluation.

That split is a reminder that "best" depends on the test, agent harness, reasoning setting and task. A composite coding-agent score measures a different collection of behaviors than a benchmark focused on resolving software issues from repositories.

Benchmarking firm Artificial Analysis reported similar trade-offs. Its GPT-5.6 review placed Sol at 59 on its Intelligence Index, one point behind Claude Fable 5, while estimating Sol's cost per task at about one-third of Fable's. It placed Sol first on its Coding Agent Index at 80.

[Click on image for larger view.] No. 2 in Artificial Analysis Intelligence Index (source: Artificial Analysis). "GPT-5.6 Sol (max) uses fewer output tokens than most models of comparable intelligence, and defines a new Pareto frontier of Intelligence vs Output Tokens per Task," the firm said. "GPT-5.6 Sol (max) offers a slight improvement in token efficiency with 15k tokens per Intelligence Index task, vs GPT-5.5 at 16k. Notably, it uses fewer tokens and is more intelligent than Claude Opus 4.8 (max), GLM-5.2 (max), and Gemini 3.5 Flash (high)."

Artificial Analysis also reported that Sol used about 15,000 output tokens per Intelligence Index task, compared with about 16,000 for GPT-5.5. It found Sol more intelligent while using fewer output tokens than Claude Opus 4.8, GLM-5.2 and Gemini 3.5 Flash at the tested settings.

The comparison comes with qualifications. Artificial Analysis said it supported OpenAI's pre-release evaluation of the GPT-5.6 family. Its results pair models with particular agent harnesses, including Sol running through Codex and Anthropic models running through Claude Code. The scores therefore measure complete tested configurations rather than model weights in isolation.

Efficiency Becomes a Product Feature

The available results support a narrower conclusion than a universal model ranking. Sol performed strongly across coding-agent, reasoning and professional-work evaluations, and several tests found favorable cost or token use relative to models with similar scores. It also trailed competing models on some prominent evaluations.

The Aug. 6 ChatGPT change brings those efficiency considerations into the user interface. Plus and Pro users no longer have to select a separate everyday model and reasoning model for each type of conversation. They can use Sol across both and vary the effort setting.

Free and Go users will receive Luna as their everyday default, reflecting OpenAI's positioning of that model as the fastest and most cost-efficient member of the GPT-5.6 family. OpenAI is consequently using different models and reasoning settings to expand access while controlling how much computation is applied to different categories of work.

For developers and enterprise administrators, that approach parallels the cost controls appearing elsewhere in the Microsoft-oriented AI ecosystem. The relevant question is increasingly not how many prompts, tokens or agent sessions a user can generate. It is how much useful work the system completes for the time, tokens and money consumed.

About the Author

David Ramel is an editor and writer at Converge 360.

PRINTABLE FORMAT

comments powered by Disqus Featured Bigger Copilot Bills? GitHub Now Lets You Peek into Per-Model Token Usage

GitHub has added input, output and cache token counts to its downloadable AI usage reports, providing more detail about how individual models consume AI Credits.

Human Skills in an AI World: How Technologists Stay Essential

Angela Dugan explains why judgment, curiosity, shared context and leadership become more valuable as AI accelerates software delivery, and previews her Live! 360 Tech Con session on using AI as a partner and career accelerator rather than a replacement.

VS Code 1.133 Flexes Claude Sessions

Microsoft's latest VS Code update adds provider switching and a GitHub-free entry path specifically for Claude sessions.

Tunability: Microsoft, GitHub Put More Dials on AI Dev Tools

New controls for model reasoning and Copilot code-review depth let developers decide how much AI effort a task warrants, with speed, depth and credit consumption all part of the tradeoff.

Subscribe on YouTube

.NET Insight Sign up for our newsletter. Email Address* Country* United States of America Afghanistan Åland Islands Albania Algeria American Samoa Andorra Angola Anguilla Antarctica Antigua and Barbuda Argentina Armenia Aruba Australia Azerbaijan Austria Bahamas Bahrain Bangladesh Barbados Belarus Belgium Belize Benin Bermuda Bhutan Bolivia, Plurinational State of Bonaire, Sint Eustatius and Saba Bosnia and Herzegovina Botswana Bouvet Island Brazil British Indian Ocean Territory Brunei Darussalam Bulgaria Burkina Faso Burundi Cambodia Cameroon Canada Cape Verde (Cabo Verde) Cayman Islands Curaçao Central African Republic Chad Chile China Christmas Island Cocos (Keeling) Islands Colombia Comoros Congo Congo, the Democratic Republic of the Cook Islands Costa Rica Côte d'Ivoire Croatia Cuba Cyprus Czech Republic Denmark Djibouti Dominica Dominican Republic Ecuador Egypt El Salvador Equatorial Guinea Eritrea Estonia Ethiopia Falkland Islands (Malvinas) Faroe Islands Fiji Finland France French Guiana French Polynesia French Southern Territories Gabon Gambia Georgia Germany Ghana Gibraltar Greece Greenland Grenada Guadeloupe Guam Guatemala Guernsey Guinea Guinea-Bissau Guyana Haiti Heard Island and McDonald Islands Holy See (Vatican City State) Honduras Hong Kong Hungary Iceland India Indonesia Iran, Islamic Republic of Iraq Ireland Isle of Man Israel Italy Jamaica Japan Jersey Jordan Kazakhstan Kenya Kiribati Korea, Democratic People's Republic of Korea, Republic of Kuwait Kyrgyzstan Lao People's Democratic Republic Latvia Lebanon Lesotho Liberia Libya Liechtenstein Lithuania Luxembourg Macao Macedonia, the former Yugoslav Republic of Madagascar Malawi Malaysia Maldives Mali Malta Marshall Islands Martinique Mauritania Mauritius Mayotte Mexico Micronesia, Federated States of Moldova, Republic of Monaco Mongolia Montenegro Montserrat Morocco Mozambique Myanmar Namibia Nauru Nepal Netherlands New Caledonia New Zealand Nicaragua Niger Nigeria Niue Norfolk Island Northern Mariana Islands Norway Pakistan Oman Palau Palestinian Territory, Occupied Panama Paraguay Papua New Guinea Peru Philippines Pitcairn Poland Portugal Puerto Rico Qatar Réunion Romania Russian Federation Rwanda Saint Barthélemy Saint Helena, Ascension and Tristan da Cunha Saint Kitts and Nevis Saint Lucia Saint Martin (French part) Saint Pierre and Miquelon Saint Vincent and the Grenadines Samoa San Marino Sao Tome and Principe Saudi Arabia Senegal Serbia Seychelles Sierra Leone Singapore Sint Maarten (Dutch part) Slovakia Slovenia Solomon Islands Somalia South Africa South Georgia and the South Sandwich Islands South Sudan Spain Sri Lanka Sudan Suriname Svalbard and Jan Mayen Eswatini (Swaziland) Sweden Switzerland Syrian Arab Republic Taiwan, Province of China Tajikistan Tanzania, United Republic of Thailand Timor-Leste Togo Tokelau Tonga Trinidad and Tobago Tunisia Turkey Turkmenistan Turks and Caicos Islands Tuvalu Uganda Ukraine United Arab Emirates United Kingdom United States Minor Outlying Islands Uruguay Uzbekistan Vanuatu Viet Nam Venezuela, Bolivarian Republic of Virgin Islands, British Virgin Islands, U.S. Wallis and Futuna Western Sahara Yemen Zambia Zimbabwe I agree to this site's Privacy Policy Please type the letters/numbers you see above. Most Popular Articles VS Code 1.133 Flexes Claude Sessions GPT-5.6 Sol Ascends for Token Efficiency; How Does It Stack Up Against Other Models? Top 5 Local AI Tools for VS Code -- All Powered by Ollama Copilot Credit Complaints Keep Coming: 'Too Expensive to Use' Tunability: Microsoft, GitHub Put More Dials on AI Dev Tools Upcoming Training Events Visual Studio Live! @ San Diego September 14-18, 2026 Live! 360 6-Week Training & Certification Course: Mastering the Microsoft AI Framework: Building Enterprise-Ready AI Agents with Microsoft Foundry October 6–November 10, 2026 VSLive! 6-Week Training & Certification Course: Blazor Developer Accelerator: Hands-On Skills for Real-World .NET Teams October 7 – November 11, 2026 Live! 360 Orlando November 15-20, 2026 Artificial Intelligence Live! Orlando November 15-20, 2026 AI Enterprise Architecture Live! Orlando November 15-20, 2026 Cybersecurity & Ransomware Live! Orlando November 15-20, 2026 Data Platform Live! Orlando November 15-20, 2026 Visual Studio Live! Orlando November 15-20, 2026 Live! 360 2-Day Hands-On Seminar: AI-Powered .NET Development with Claude & Claude Code December 8-9, 2026 VSLive! 4-Day Hands-On Training Seminar: Immersive .NET Full Stack Training with CoPilot: 4-Day Hands-On Experience December 15-18, 2026 Visual Studio Live! Las Vegas March 22-26, 2027 Visual Studio Live! @ Microsoft HQ August 2-6, 2027 Free Webcasts Operating at galactic scale: Inside Sumo Logic’s agentic security program Called it (mostly): Checking in on 2026 predictions so far 2026 Security operations insights Build and Automate Your Infrastructure with Red Hat on Microsoft Azure

More Webcasts

CONTACT USADVERTISETRAININGFREE NEWSLETTERSSITE MAPREPRINTSLIST RENTALGLOSSARY

AI Boardroom ADTmag AWS Insider Campus Security Today Campus Technology Environmental Protection Live! 360 Events MCPmag MedCloudInsider Occupational Health & Safety Pure AI Redmond Redmond Channel Partner Security Today Spaces 4 Learning TechMentor Tech Tactics in Education The AI Pivot THE Journal Virtualization & Cloud Review Visual Studio Live!

©1996-2026 1105 Media Inc. See our Privacy Policy, Cookie Policy and Terms of Use. CA: Do Not Sell My Personal Info

Problems? Questions? Feedback? E-mail us.

What this supports

  • Supports describing the Plus/Pro slider and separating product behavior from the earlier benchmark.
  • Supports the narrower conclusion that Sol is strong but not best everywhere.

What this does not support

  • Does not support converting a ChatGPT subscription into API billing or generalizing error reduction to all questions.
  • The August product update and July benchmark are different snapshots.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Visual Studio Magazine · David Ramel · Original publication date 2026-08-06 · Site edit date 2026-09-20

Open original source

GPT-5.6 Sol

Compare GPT-5.6 Sol in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Sol: Specs, Access, Changes, and the Risks That Still Matter

OpenAI's current GPT-5.6 Sol model page lists a 1.05M context window, 128K max output, reasoning controls, and a time-sensitive API price card. Here is what those facts mean for API, Codex, and browser users.

Related reviews

Artificial Analysis: Sol's intelligence, coding-agent result, and cost per taskArtificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.METR: Sol's time horizon changes with cheating treatmentIn Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.OpenAI release note: Sol's official results on long-horizon, coding, and knowledge workOpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.Every: Sol excels as a collaborative knowledge-work partner, not as judgmentEvery describes Sol as fast and steerable across 24 drafts, email, meetings, and retrieval, but it scored 56/100 versus Fable's 90/100 on Senior Engineer and ranked last of six in writing; collaboration is not autonomous judgment.Route ChatGPT tasks through Sol’s reasoning settingsRun the same task at faster and deeper reasoning settings: prefer speed for short questions, then increase reasoning for planning, research, writing, coding, and decisions; OpenAI’s 68% figure is an internal relative change, not public accuracy.Design a verifiable multi-agent workflow with the Responses APISeparate judgment from deterministic processing, then combine programmatic tool calls, parallel subagents, and prompt-cache boundaries into a long-running workflow whose cost, latency, citations, and failures can be reviewed.Configure Codex for a million-token context and auto-compactionThe source shows config.toml and one-session CLI examples for the model ID, a 1,000,000-token context budget, and a 900,000-token compaction threshold; confirm client support and keep a rollback configuration before editing.Deliver code with prediction, planning, review, and verificationSplit long-running coding into prediction, planning, implementation, adversarial review, and independent verification, checking the plan, tests, and stop conditions item by item; this is a commenter’s personal workflow, not Codex’s default configuration.