Design a verifiable multi-agent workflow with the Responses API
Separate judgment from deterministic processing, then combine programmatic tool calls, parallel subagents, and prompt-cache boundaries into a long-running workflow whose cost, latency, citations, and failures can be reviewed.
Prepare
user goal, tool schemas, agent responsibilities, cache boundaries, acceptance criteria
Runtime
OpenAI Responses API and a production agent harness; define tool calls and human checkpoints yourself
Route ChatGPT tasks through Sol’s reasoning settings
Run the same task at faster and deeper reasoning settings: prefer speed for short questions, then increase reasoning for planning, research, writing, coding, and decisions; OpenAI’s 68% figure is an internal relative change, not public accuracy.
Prepare
task type, fixed acceptance criteria, speed and quality preference, client and date record
Runtime
ChatGPT web, mobile, or desktop client; do not infer Codex or API parameters
Configure Codex for a million-token context and auto-compaction
The source shows config.toml and one-session CLI examples for the model ID, a 1,000,000-token context budget, and a 900,000-token compaction threshold; confirm client support and keep a rollback configuration before editing.
Deliver code with prediction, planning, review, and verification
Split long-running coding into prediction, planning, implementation, adversarial review, and independent verification, checking the plan, tests, and stop conditions item by item; this is a commenter’s personal workflow, not Codex’s default configuration.
OpenAI release note: Sol's official results on long-horizon, coding, and knowledge work
OpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.
Evidence
Vendor report
Boundary
Does not support independent reproduction, a universal cross-model ranking, or current product availability.
Artificial Analysis: Sol's intelligence, coding-agent result, and cost per task
Artificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.
Evidence
Independent measurement
Boundary
Does not support extrapolating 59, 80, or cost per task to other clients, live prices, or all tasks.
CodeRabbit: Sol's trade-offs in long coding-agent runs and code review
CodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.
Evidence
Independent measurement
Boundary
Does not support general success or precision rates independent of CodeRabbit's harness.