GPT-5.6 Sol · workflow
Split long-running coding into prediction, planning, implementation, adversarial review, and independent verification, checking the plan, tests, and stop conditions item by item; this is a commenter’s personal workflow, not Codex’s default configuration.
Goal: {{GOAL}}
Repository: {{REPOSITORY}}
Record expected behavior and risks; plan the scope, immutable interfaces, and stop conditions; implement only the current plan and run {{TEST_COMMAND}}; adversarially review the code and tests; submit test, human-review, and rollback evidence. On verification failure, fix only the recorded gap and do not expand scope after the stop conditions are met.Replace before running: {{GOAL}}, {{REPOSITORY}}, {{TEST_COMMAND}}
Produce a staged coding record with stop conditions: write expectations and risks, plan the scope, implement while testing, then use independent review and verification to distinguish “the model says it is done” from “the artifact meets the requirements.” This is an editorial adaptation of a commenter’s personal hack.
The coding agent can save plans, read the repository, run targeted tests, and inspect diffs.
Define the goal, interfaces that must not change, risks, test commands, human approval points, and stop conditions.
Prepare independent checks for every acceptance signal, especially for permissions, databases, and production deployment changes.
Prediction: record expected behavior, unknowns, key risks, and minimum acceptance signals.
Planning: break the goal into checkable steps and lock scope, interfaces, file boundaries, rollback, and stop conditions.
Implementation: execute only the current plan and run relevant tests as changes are made; do not expand for hypothetical future needs.
Review: adversarially inspect code and tests for reward hacking, overengineering, weakened assertions, and missed regressions.
Verification: use independent commands, black-box checks, or human review to compare the artifact with the plan and acceptance signals; return only to the missing stage when it fails.
Loop: fix only recorded failures; pause for human scope confirmation when the same area changes repeatedly or a high-risk boundary is reached.
Prediction, plan, implementation diff, test results, review findings, and verifier decision are recorded.
Tests cover the requirements; assertions were not removed, weakened, or scope-shifted to pass.
Every verification item is marked pass, fail, or not run with command, log, or human-review evidence.
Failures trigger only necessary local fixes; no scope expansion occurs after the stop conditions are met.
Pause and reconfirm goal, scope, and acceptance criteria when the plan causes goal drift.
Record the smallest test gap and fix the implementation; never weaken a test to hide a failure.
Stop the loop and ask for human judgment when the agent repeats changes, adds abstractions, or exhausts its quota.
If independent verification cannot run, state the blocker and next check instead of claiming success.
The source provides no complete public configuration or controlled benchmark; map stage names to the actual agent hooks, files, or commands.
This is not Sol’s or Codex’s default behavior and does not establish success rate, cost, or production readiness.
Verification cannot replace human review, permission approval, or release controls for high-risk changes.
Breaking Sol's long-running coding tasks into five stages—prediction, planning, implementation, adversarial review, and verification—makes it possible to separately check whether the “model claims completion” and whether the work is “delivered according to plan.”
Suitable tasks: Multi-file coding, agent orchestration, test-driven implementation, and long-running tasks prone to endless repair loops.
Unsuitable tasks: One-off Q&A or creative tasks with no verifiable output.
Applicable model version: GPT‑5.6 Sol; it can also be transferred to other coding agents.
Applicable clients, Agents, or APIs: Agents such as Codex and OpenCode that can save plans, run tests, and review code.
Recommended reasoning tier and parameters: The comment did not specify a fixed tier; start with the lowest tier that can complete the task, and track costs separately for iterative tasks.
The following is a reusable skeleton reconstructed from the stage names made public in the comment, not the author's complete public system prompt:
prediction_stage:
Write down expected behavior, key risks, and the minimum acceptance signals.
planning_stage:
Break the goal into checkable steps; define scope, interfaces that must not change, and stop conditions.
implementation_stage:
Implement only the current plan, running targeted tests as you go.
review_stage:
Adversarially inspect the implementation and tests, focusing on reward hacking, overengineering, and tests that go easy on the implementation.
verifier_stage:
Compare the actual output item by item against the plan and acceptance signals; deliver if it passes, and return only to the missing stage if it fails.
loop_policy:
If verification fails, state the specific gap and next step; once the stop conditions are met, prohibit further scope expansion.Save the prediction and plan before starting the task to prevent goal drift later.
During implementation, submit only code and tests related to the current plan.
During review, check whether the tests truly cover the requirements and whether assertions were weakened just to “pass.”
During verification, run independent commands or black-box checks and record pass/fail item by item.
In a loop, fix only failed items; if the model repeatedly changes the same area, pause and ask a human to confirm the scope.
The sequence publicly mentioned in the comment was prediction_stage, planning_stage, review_stage, and verifier_stage, with a recommendation to check for reward hacking after the code and tests are in place.
The commenter explicitly said this was a hack from their own testing, not a controlled benchmark or an official Codex configuration.
Experiences of Sol Ultra/Max in the main post of the same thread are strongly disputed; some users in the comments reported efficiency, while others reported overengineering and exhausting their quotas.
This is a community workflow suggestion, not Sol's default behavior or an official OpenAI best practice.
The stage names and examples need to be mapped to the actual Agent's hooks, files, or commands; the source does not provide a complete configuration file that can be copied.
The verification stage cannot replace human review of high-risk changes, especially database, permission, and production deployment operations.
The comment recommends using review_stage after implementation and testing are complete to check for “signs reward hacking in tests”.
Reddit r/codex · Source date: 2026-07-23 · Edited: 2026-09-20
Read the original sourceGPT-5.6 Sol
Run this guide in the environment listed above. Downloading does not transfer the template or establish model availability for your account.