Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Sol · Community source · Independent measurement

Box Complex Work: Sol's advantage is concentrated in quantitative document chains

In Box's twelve-industry document evaluation, Sol scores 76%, 74%, 72%, 61%, 60%, and 58% in six named industries versus GPT-5.5 at 71%, 63%, 66%, 54%, 51%, and 46%; the full sample and scorer are undisclosed.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceIndependent measurementEdited 2026-09-20

Test conditions

Model version
GPT-5.6 Sol; compared with Terra, Luna, and GPT-5.5
Provider / client
Box Complex Work Eval; harness not disclosed
Reasoning tier
Not disclosed
Tools
Document reading, number checking, diligence, and quantitative analysis; permissions not disclosed
Task set
Twelve-industry document tasks including forecasts, grade recalculation, retail, energy, life sciences, and clinical work
Sample / repeats
Full sample, scorer, seeds, and repeats not disclosed
Publication / collection date
2026-07-10 / 2026-08-18
Traceable results
Six-industry Sol/GPT-5.5: 76/71, 74/63, 72/66, 61/54, 60/51, 58/46%; data analysis 64/57%

Key data and applicable tasks

Test environment

  • Benchmark: Box Complex Work Eval, covering real document-driven tasks across twelve industries.

  • Task types: Reading source documents, checking numbers, due diligence, identifying errors in expert outputs, and quantitative analysis.

  • Comparison: GPT‑5.6 Sol, Terra, Luna, and GPT‑5.5; the page does not disclose the complete sample count or harness version.

Inputs/configuration

  • Inputs consisted of document tasks such as multi-year financial forecasts, student-grade recalculations, retail product performance, energy operations reports, life-sciences test records, and clinical cases.

  • Key acceptance criteria were whether the chain of numbers remained consistent, the denominator was correct, reports were evidence-based, and high-risk judgment errors were avoided.

  • Box positions the model as an enterprise workflow option for future Box AI / Box AI Studio offerings.

Results data

Industry taskGPT‑5.6 SolGPT‑5.5
Financial Services76%71%
Public Sector74%63%
Retail72%66%
Energy61%54%
Life Sciences60%51%
Healthcare58%46%
  • Overall, Sol scored about 1% higher than GPT‑5.5, but Box believes the gap is concentrated in more difficult, higher-consequence quantitative tasks.

  • Across data-analysis tasks, Sol scored 64% versus GPT‑5.5's 57%; the breakdown was 80% versus 63% in retail, 51% versus 27% in life sciences, 62% versus 50% in healthcare, and 78% versus 75% in finance.

  • For average latency, Terra was about 16% faster than Sol and Luna about 19% faster; in retail inventory/profit analysis, Luna returned a complete answer in roughly half Sol's time.

Conclusions

Sol's advantage looks more like getting the chain of numbers in source documents right and keeping it defensible, rather than leading uniformly across all knowledge work. Try Sol first for high-stakes quantitative analysis; route repetitive, structured, high-throughput tasks to Terra/Luna, using accuracy and latency together.

Limitations

  • Box's internal benchmark, graders, sample size, random seeds, and original documents were not made public, so outsiders cannot fully rerun it.

  • Box is a potential product partner, creating publication-selection and scenario bias; the results cannot replace blind cross-vendor testing.

  • The page describes the overall improvement as “about 1%” but does not provide the total sample count or confidence interval.

Reproduction steps

  1. Build a task set spanning twelve industries, retaining source documents, target answers, and numeric-checking rules.

  2. Fix the context, tools, and output format for Sol, Terra, Luna, and GPT‑5.5.

  3. Design separate graders for financial forecasts, denominator calculations, record recalculations, and clinical judgments.

  4. Record task accuracy, latency, cost, and error types; do not record only the overall average score.

  5. Conduct external validation with hidden documents and a second batch of industry tasks.

Original evidence and data

  • Box explicitly framed “reading source files, checking numbers, running due diligence, and reviewing expert outputs” as benchmark tasks, rather than a questionnaire about chat preferences.

  • The article gives Sol/GPT‑5.5 scores for six industries and scores for the data-analysis subset.

Scope of applicability

  • The conclusion is better suited to document-driven enterprise analysis; it does not mean Sol is likewise ahead in open-ended writing, frontend visuals, or long-horizon coding.

  • Latency is reported only as relative percentages, without service-region, concurrency, or token statistics.

  • This is Box's own evaluation and should not be combined directly with Artificial Analysis scores into a unified ranking.

Source excerpt or observation (compliance short quote only)

Box summarized the overall result by saying that Sol “edges past” GPT‑5.5, but then emphasized that the advantage was concentrated in the most difficult quantitative tasks.

What this supports

  • Supports limiting the reported advantage to document-driven, numerically consistent, high-consequence analysis.
  • Supports viewing industry results instead of only the roughly 1% overall description.

What this does not support

  • Does not support equivalent conclusions for open-ended writing, visual builds, or coding.
  • Box is a potential product partner; its values must not be combined with AA into one ranking.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X (Box) · @sidharths00 / Box · Original publication date 2026-07-10 · Site edit date 2026-09-20

Open original source

GPT-5.6 Sol

Compare GPT-5.6 Sol in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Sol: Specs, Access, Changes, and the Risks That Still Matter

OpenAI's current GPT-5.6 Sol model page lists a 1.05M context window, 128K max output, reasoning controls, and a time-sensitive API price card. Here is what those facts mean for API, Codex, and browser users.

Related reviews

METR: Sol's time horizon changes with cheating treatmentIn Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.OpenAI system card: Sol's tool-safety scores and confirmation boundariesThe system card reports Sol at 1.000 on connector injection and 0.910 on search/function-call cases, with computer-use confirmation scores of 0.98/0.99/0.93 for financial, high-risk, and general confirmation; these are safety evaluations, not task success rates.OpenAI release note: Sol's official results on long-horizon, coding, and knowledge workOpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.Artificial Analysis: Sol's intelligence, coding-agent result, and cost per taskArtificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.Structure Sol requests for long documents, images, and retrievalUse role, goal, format, and criteria; map a long document before asking targeted questions, state exactly what an image should yield, and request sources for fresh facts. Client-specific capability claims still require verification at the actual entry point.Deliver code with prediction, planning, review, and verificationSplit long-running coding into prediction, planning, implementation, adversarial review, and independent verification, checking the plan, tests, and stop conditions item by item; this is a commenter’s personal workflow, not Codex’s default configuration.Route ChatGPT tasks through Sol’s reasoning settingsRun the same task at faster and deeper reasoning settings: prefer speed for short questions, then increase reasoning for planning, research, writing, coding, and decisions; OpenAI’s 68% figure is an internal relative change, not public accuracy.Choose prompt scope and reasoning effort for SolStart with the smallest prompt and an effort level suited to the task, then test whether Pro mode, tool calling, or extra constraints improve a representative workload; the article also compares Claude Fable 5 and is not an independent OpenAI specification.