Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Sol · Community source · Independent measurement

SlopCodeBench: Sol ties Fable on strict long-horizon coding passes

SlopCodeBench tests long-horizon coding with six challenges, 30 incremental checkpoints, and fresh contexts; Sol in Codex CLI 0.145.0 passed 10/30 strict checkpoints (33.3%), tying Fable in one run per model.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceIndependent measurementEdited 2026-09-20

Test conditions

Model version
GPT-5.6 Sol; compared with Fable 5 and two Kimi K3 providers
Provider / client
Sol Codex CLI 0.145.0; Fable Claude Code 2.1.219; Kimi OpenCode 1.18.0
Reasoning tier
Not disclosed
Tools
Corresponding CLI/coding agent; fresh context per checkpoint
Task set
Six challenges—xjq, file_backup, dag_execution, circuit_eval, code_search, etl_pipeline—across 30 checkpoints
Sample / repeats
One run per model/provider; Sol 10/30; repeats and seeds not disclosed
Publication / collection date
2026-08-05 / 2026-08-18
Traceable results
Sol/Fable 10/30 (33.3%); Kimi 8/30 and 7/30; Sol had 14 isolated passes

Key data and applicable tasks

Test environment

  • Benchmark: SlopCodeBench. The tasks do not reveal all requirements at once; they add requirements progressively across multiple checkpoints to test the risk of codebase degradation over time.

  • Scale: 6 challenges and 30 checkpoints, with a fresh context for each checkpoint; each of the four models was run once.

  • Harness: Fable—Claude Code 2.1.219; Sol—Codex CLI 0.145.0; Kimi K3 (Baseten/Modal)—OpenCode 1.18.0.

  • Scoring: strict pass; a checkpoint counts as passed only when the current checkpoint and all black-box tests inherited from previous checkpoints pass.

Inputs/configuration

  • xjq: 5 checkpoints, expanding from XPath/XML/HTML/JSON queries to selectors, structured output, and union.

  • file_backup: 4 checkpoints, expanding YAML backup scheduling to archive, verify, targets, and incremental state.

  • dag_execution: 3 checkpoints covering a task-pipeline DSL, caching, and dynamic cache coverage.

  • circuit_eval: 8 checkpoints covering circuit parsing, vectors, three-valued logic, formatting, statistics, equivalence, and optimization.

  • code_search: 5 checkpoints covering multilingual text/regex/structured search, AST selectors, and fixes.

  • etl_pipeline: 5 checkpoints covering JSON ETL, branching, definitions, parameterized calls, and namespace composition.

  • The author publicly released condensed requirements for all 30 checkpoints, the model order, isolated contexts, and the black-box judging mechanism.

Results

ModelStrict checkpoint pass rateStrict passes
Fable 533.3%10/30
GPT‑5.6 Sol33.3%10/30
Kimi K3 (Modal)26.7%8/30
Kimi K3 (Baseten)23.3%7/30
  • As a tiebreaker, Fable had 16 isolated passes, compared with Sol's 14.

  • In the final snapshot, approximately 95% of Sol's code lines triggered at least one slop rule, compared with 86% for Fable and 82%/79% for Kimi K3; the author believes the rules may be too aggressive.

  • Sol left 1,318 SLOC in persistent Python test files; the other agents used shell, fixtures, or temporary files, so the figures are not directly comparable.

  • The author stresses that these are directional results from one run per model/provider and should not be treated as statistically significant rankings.

Conclusion

On long-horizon coding tasks with progressively disclosed requirements and inherited defects, Sol and Fable had the same strict pass rate, but all models accumulated defects as checkpoints progressed. Sol's code-volume/test-artifact metrics also suggest that “more complete testing” does not equal lower complexity; black-box regression and complexity audits should be run together.

Limitations

  • Each model and Kimi provider was run only once, leaving too small a sample for significance inferences.

  • Sol, Fable, and Kimi used different harnesses; differences in cost and tooling affect the results.

  • The relationship between the slop metric and “code maintainability” has not been validated, and the author also notes that the rules may penalize too heavily.

  • The article does not disclose all original repository snapshots for each checkpoint or complete token/latency logs.

Reproduction steps

  1. Obtain the SlopCodeBench paper, runner, and problem catalog, and pin the versions of all six challenges.

  2. Create a clean context for each checkpoint and expose only the current requirements; retain the code from the previous checkpoint.

  3. Pin each model's harness version and tool permissions, and record every code submission, test output, and tool call.

  4. Run hidden black-box tests for the current checkpoint and all inherited checkpoints, and score strict passes.

  5. Also measure defect accumulation, code growth, duplication, cognitive complexity, and test-file types.

  6. Repeat multiple times and cross-run different harnesses before discussing statistical differences.

Original evidence and data

  • The article provides the task shapes for 6 challenges, condensed inputs for 30 checkpoints, harness versions, and strict scoring rules.

  • The strict pass counts for Sol/Fable/Kimi and some code-quality statistics are disclosed directly, and the author explicitly states the sponsorship and one-run limitations.

Scope of applicability

  • Best suited for evaluating long-horizon coding, progressively added requirements, and regression-defect propagation; it is not equivalent to ordinary single-turn code generation.

  • The 33.3% tie between Sol and Fable does not imply that they are equivalent in creativity, writing, or knowledge work.

  • The “slop” results should be treated only as signals requiring validation, not as direct conclusions about engineering quality.

Source excerpt or observation (short quote for compliance only)

The author characterizes the experiment as “directional, not exhaustive” and emphasizes that all new models continued to accumulate defects.

What this supports

  • Supports evaluating incremental requirements, defect propagation, and long-horizon coding rather than one-turn generation.
  • Supports treating the 33.3% Sol/Fable tie as one directional result.

What this does not support

  • Does not support a statistically significant ranking or equivalence in creative, writing, or knowledge work.
  • Different harnesses, one run, and disputed slop rules affect the result.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X · dexhorthy · Original publication date 2026-08-05 · Site edit date 2026-09-20

Open original source

GPT-5.6 Sol

Compare GPT-5.6 Sol in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Sol: Specs, Access, Changes, and the Risks That Still Matter

OpenAI's current GPT-5.6 Sol model page lists a 1.05M context window, 128K max output, reasoning controls, and a time-sensitive API price card. Here is what those facts mean for API, Codex, and browser users.

Related reviews

CodeRabbit: Sol's trade-offs in long coding-agent runs and code reviewCodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.METR: Sol's time horizon changes with cheating treatmentIn Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.Reddit Cursor: one backend implementation comparison with Sol mediumA Reddit user ran Grok 4.6 extra high and Sol medium in Cursor on the same roughly 2,500-line backend plan, with Fable 5 high as judge; the author gives Sol an approximate 60/40 subjective win in one run.Lynkr ITSMBench: routing lowers cost while Sol's binary pass rate remains limitedLynkr routed Sol through pi on 89 enterprise IT-service tasks: 31% full-suite Pass@1, 35%/40% matched Pass@1/Pass@2, about $0.87–$0.90 per task, and 92–95% cache hits; many failures missed only a few assertions.Deliver code with prediction, planning, review, and verificationSplit long-running coding into prediction, planning, implementation, adversarial review, and independent verification, checking the plan, tests, and stop conditions item by item; this is a commenter’s personal workflow, not Codex’s default configuration.Configure Codex for a million-token context and auto-compactionThe source shows config.toml and one-session CLI examples for the model ID, a 1,000,000-token context budget, and a 900,000-token compaction threshold; confirm client support and keep a rollback configuration before editing.Give Codex an Occam rule against over-engineeringAsk a coding agent to choose the simplest implementation that satisfies demonstrated requirements, reuse or remove existing code before adding layers, and keep clear module boundaries; the rule is community guidance, not a guarantee.Design a verifiable multi-agent workflow with the Responses APISeparate judgment from deterministic processing, then combine programmatic tool calls, parallel subagents, and prompt-cache boundaries into a long-running workflow whose cost, latency, citations, and failures can be reviewed.