Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Sol · Community source · Personal experience

Nate Herk: Sol costs less on creative builds, while Fable wins more blind selections

Nate Herk used the same /goal to compare Sol in Codex with Fable in Claude Code: Sol won the roughly seven-minute/$1 visual-object build, while Fable was selected for the bike game and scrolling site; refusals confounded the small API sample.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model version
GPT-5.6 Sol; compared with Claude Fable 5
Provider / client
Sol used Codex; Fable used Claude Code
Reasoning tier
Not disclosed
Tools
Agent /goal builds and stateless API rounds; permissions not disclosed
Task set
Bike game, scrolling site, five visual objects, and API rounds
Sample / repeats
Each build once; API recorded 24 Sol and 3 Fable responses; author blind-selected
Publication / collection date
2026-07-10 / 2026-08-18
Traceable results
Sol: bike about 23m/$4.50, site about 7m/$1, objects about 7m/$1; Fable won first two

Key data and applicable tasks

Test environment

  • Agent build: The same /goal prompt with complete creative freedom; Fable ran in Claude Code, Sol in Codex; the author reviewed the results blind before revealing the models.

  • Three builds: A playable bicycle game, an interactive scrolling website, and five completely different visual objects.

  • API runs: Fast, stateless tasks without an Agent loop, comparing response rate, capability score, speed, and cost.

Inputs/configuration

  • The bicycle-game prompt called for an open-world game in the browser, with WASD steering, spacebar jumping, Q/E aerial tricks, and Shift acceleration.

  • The website prompt asked for “the most impressive interactive scrolling website.”

  • The third task only asked for five fundamentally different visual elements and required the model to return a gallery and five sister sites.

  • These builds used different harnesses, so the comparison was of complete working configurations rather than bare models.

Results data

TaskFable 5GPT‑5.6 SolAuthor's choice
Bicycle game21m37s, $14.22, about 90k output tokens23m, $4.50, about 31kFable
Interactive scrolling website23m, $19.24, about 80kabout 7m, about $1, about 20kFable
Five visual objects15m, about $15, about 65k7m, about $1, about 22kSol
  • The API fast-task response records were 24 for Sol and 3 for Fable; the author noted that most of the difference came from Fable refusing to answer.

  • Among the answers that were actually returned, Sol's capability score was 0.98 and Fable's was 0.966; the batch cost $16 for Sol and $63 for Fable.

  • The author's routing judgment was Fable for management/strategy, and Sol for execution, verification, and delivery.

Conclusions

In this personal blind test, Sol was clearly more efficient in tokens and cost, and won on the open-ended “five objects” task; Fable was more often chosen for the final aesthetics and completeness of creative builds. The result supports the “Fable sets direction, Sol executes” workflow hypothesis, but it is not an overall ranking of model capabilities.

Limitations

  • Each of the three builds was run once, so subjective choices and the author's design preferences affect the results.

  • The toolchains, default prompts, and context management in Codex and Claude Code differed; the cost difference cannot all be attributed to the models.

  • The 24–3 API gap is confounded by refusals; 0.98 and 0.966 are not public standard benchmarks either.

  • The article does not provide the complete API inputs, grader, latency distribution, or random seed.

Reproduction steps

  1. In two isolated repositories, fix the same commit, identical /goal text, and identical asset permissions.

  2. Randomize model run order and clear git history and caches, preventing the second model from reading the first model's artifacts.

  3. For each build, record completion time, input/output tokens, tool calls, playability, and blind-review results.

  4. For each API task, record “complete, refusal, error, or timeout”; do not conflate refusals with capability failures.

  5. Repeat for multiple rounds and report the mean, dispersion, and harness differences.

Original evidence and data

  • The article publicly disclosed the intended prompts, time, cost, and approximate output tokens for the three builds.

  • The author explicitly said Sol cost about half as much in tokens as Fable, and considered Sol's price tier closer to Opus 4.8 for comparison.

Scope of applicability

  • Suitable for considering cost/quality trade-offs in Agent builds, but not as a general benchmark for writing, mathematics, or coding.

  • “Sol is the worker” is the author's empirical model; it could reverse under other harnesses, task boundaries, or design standards.

  • For creative outputs, define the blind-review rubric in advance to avoid presenting personal aesthetics as an objective win/loss.

Source excerpt or observation (compliance short quote only)

The author's core routing metaphor was “Fable is the manager, Sol is the worker”.

What this supports

  • Supports separating cost, speed, and final selection in complete agent workflows.
  • Supports the conditional personal routing hypothesis of Fable for direction and Sol for execution.

What this does not support

  • Does not support a bare-model cost comparison or universal creative ranking; the harnesses differ.
  • Three builds and the tiny API sample cannot estimate a stable win rate, and refusals changed the sample.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X · Nate Herk (@nateherk) · Original publication date 2026-07-10 · Site edit date 2026-09-20

Open original source

GPT-5.6 Sol

Compare GPT-5.6 Sol in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Sol: Specs, Access, Changes, and the Risks That Still Matter

OpenAI's current GPT-5.6 Sol model page lists a 1.05M context window, 128K max output, reasoning controls, and a time-sensitive API price card. Here is what those facts mean for API, Codex, and browser users.

Related reviews

Lynkr ITSMBench: routing lowers cost while Sol's binary pass rate remains limitedLynkr routed Sol through pi on 89 enterprise IT-service tasks: 31% full-suite Pass@1, 35%/40% matched Pass@1/Pass@2, about $0.87–$0.90 per task, and 92–95% cache hits; many failures missed only a few assertions.X one-shot visual build: Sol looked better, but the game was broken and cost moreIn a Command Code comparison with the same /design prompt and one attempt, Sol cost $0.32 and produced a nice interface but an unplayable browser game; GLM 5.3 cost $0.016 and was playable, while Opus 5 cost $0.37 and had the strongest clone.OpenAI release note: Sol's official results on long-horizon, coding, and knowledge workOpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.Artificial Analysis: Sol's intelligence, coding-agent result, and cost per taskArtificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.Configure Codex for a million-token context and auto-compactionThe source shows config.toml and one-session CLI examples for the model ID, a 1,000,000-token context budget, and a 900,000-token compaction threshold; confirm client support and keep a rollback configuration before editing.Deliver code with prediction, planning, review, and verificationSplit long-running coding into prediction, planning, implementation, adversarial review, and independent verification, checking the plan, tests, and stop conditions item by item; this is a commenter’s personal workflow, not Codex’s default configuration.Design a verifiable multi-agent workflow with the Responses APISeparate judgment from deterministic processing, then combine programmatic tool calls, parallel subagents, and prompt-cache boundaries into a long-running workflow whose cost, latency, citations, and failures can be reviewed.Give Codex an Occam rule against over-engineeringAsk a coding agent to choose the simplest implementation that satisfies demonstrated requirements, reuse or remove existing code before adding layers, and keep clear module boundaries; the rule is community guidance, not a guarantee.