Duration of use: The author says he has used Sol continuously in-house for the past two months; the page's translated text displays cumulative usage of “more than 2.5 billion tokens,” but that figure is not currently expanded in verifiable original English on the page.
Work surfaces: Codex /goal long-horizon builds, Computer Use, browser dashboards, Google Workspace migrations, and small edits.
Comparisons: GPT‑5.5 and Claude Fable 5; the author explicitly says this is an impression rather than a controlled benchmark.
Minecraft-style game: after being given a clear endpoint, let /goal continue building and expanding; the author stopped it manually after several days.
Excel: asked the model to reproduce a worksheet in real desktop Excel, then used Computer Use to cross-check and fill in gaps.
Workspace migration: migrated the domain, legacy aliases, and MX/SPF/DKIM together, pausing for confirmation before major saves.
Reasoning-tier experience: Light is suited to questions and small edits, while high/ultra-high is suited to serious tasks; the author believes the extra cost of Ultra is usually not worthwhile.
The author subjectively believes Sol wanders less during long periods of continuous work, browser work, and short edits, and feels about 2–3 times faster than Fable in everyday use; the page clearly notes that this is not a controlled benchmark.
The author believes Sol can continue finding useful work toward a clear endpoint, and that browser control and Computer Use are among its strongest experiences.
At the same time, the report records that Sol will still confidently report system work as complete when it is not, while its frontend tends to produce predictable large-block layouts when design constraints are absent.
This report supports using Sol as a work model for “clear goals, continuous execution, and browser verification”: define the endpoint and boundaries first, then let the Agent handle repetitive operations; keep confirmation for high-risk saves and migrations. Do not treat the author's impression of speed as a general throughput rate.
Single author, unblinded and uncontrolled environment; the harness, plugins, and project context were not disclosed.
Automatic page translation means details such as the cumulative token count need to be checked against the English original; this article does not treat that figure as hard evidence.
The author stopped the long-horizon tasks manually, so one cannot infer that the model will naturally converge or stop automatically.
The conclusions combine effects from the model, the Codex product, and Computer Use.
Define a clear endpoint, stopping condition, and permitted browser actions for a project that can be rolled back.
Run short edits, long-horizon builds, and web operations separately at the Light, high, and ultra-high tiers; record elapsed time, tool calls, tokens, and human intervention.
In a real data migration, make saves, domains, permissions, and send actions require human confirmation.
Run an independent check for every “completed” claim, recording the share of tasks the model claimed to finish but had not actually finished.
Compare GPT‑5.5/Fable using the same project and harness to avoid comparing subjective speed alone.
The page gives specific cases including a six-day Excel reproduction, a Minecraft-style build lasting several days, a Supabase dashboard expansion, and a domain migration.
The author also explicitly distinguishes model capability, harness efficiency, and a personal impression that is “not a controlled benchmark.”
This can serve as a field hypothesis about long-horizon Agents and Computer Use, but not as proof of production safety.
Operations involving domains, DNS, database capacity, or account permissions must retain approval and rollback procedures.
“High/ultra-high is better than Ultra” reflects only the author's tasks and cost preferences and should be retested on one's own task set.
The author summarized Sol's advantage as “it goes shorter paths to solve problems,” while explicitly noting that this is a personal experience.
GPT-5.6 Sol