This small-sample local benchmark does not prove that Terra is the best general-purpose option, but it shows Terra max scoring 4.86/5 in one primary-output role and getting 4/4 correct on a planted-bug repair while taking 102.20 seconds, making it suitable as a routing candidate rather than a default conclusion.
Tasks: 3 strategic-decision memos, 1 repository-based execution brief, and 2 planted-bug repair tasks with direct tests.
Comparisons: GPT-5.6 Sol/Luna/Terra, along with routes such as Fable, GPT-5.5, and GPT-5.4 mini.
Controls: The same prompt or fixture was used within each comparison; strategic tasks used a fixed rubric covering judgment, supporting evidence, risks, boundaries, executability, and efficiency.
Evaluation method: Model identity was visible to evaluators; the author explicitly described these as local workflow scores, not general intelligence scores.
The author published the task types, routes, observed times, reasoning tokens, and some results, but did not release the complete prompts, repository fixtures, test code, or downloadable traces. This source is suitable for reproducing the “role-based few-shot routing” method, but not enough to recompute every score.
Multi-layer productization: Fable reference 95; Sol max 94, xhigh 90, high 87, GPT-5.5 high 81.
Decision-method / wrong-object review: Fable reference 92; Sol max 91, xhigh 90, Opus-class 87.
Bounded opportunity comparison: Fable reference 95; Sol high 94, Sol max 90, Luna max 88, Terra max 88.
When Sol high was within 1 point of the reference, it took 81.5 seconds and 369 reasoning tokens; Sol max took 216.9 seconds and 5,178 tokens, without changing the decision.
Repository-grounded brief: Sol high scored 93 in 80.73 seconds with 1,818 reasoning tokens; Sol medium scored 91 in 70.66 seconds with 779 tokens. The author calculated that medium was approximately 12.5% faster and used 57% fewer reasoning tokens.
Planted-bug repair: Terra max had 1 fixture, scored 4/4 on direct tests, and passed on the first round, with an observed time of 102.20 seconds; Sol low in the same table had 2 fixtures, 10/10, and an average of approximately 37.2 seconds.
Another frozen role-specific suite: Terra max scored 4.86/5 in the primary-output role; Sol xhigh scored 4.70 and 4.93 in the quality-control and adversarial-review roles, respectively.
Terra max performed strongly in one of the author’s “primary-output” roles, but was correct and slow on the public planted-bug task, so this does not justify replacing every Sol route.
The more valuable reusable conclusion is to tier routes by task role and treat differences of “one point or less” as noise; the author recommends a blinded holdout as the next step.
A Terra rerun should prioritize first-round pass rate on the same fixture, time, reasoning tokens, and whether manual correction is needed, rather than comparing only the final text’s overall impression.
The sample is extremely small: 3 strategic memos, 1 repository brief, and 1–2 bug fixtures, leaving substantial statistical uncertainty.
Evaluators knew the model identities, creating confirmation bias; the author also framed the results as local workflow scores.
Latency from the local CLI includes environmental overhead, and subscription consumption cannot be directly equated with API token cost.
The models used different Chat, CLI, Agent harness, and multi-agent surfaces; the cross-model results were not run under fully identical configurations.
Prepare minimal but fixed task sets and acceptance rubrics for “strategic decisions, repository briefs, planted-bug repairs, and primary output.”
Run Terra max under the same working tree, tool versions, timeouts, and test commands, and save the complete prompts, outputs, tool traces, tokens, and timings.
Hide model identities from the evaluation table and have at least a second evaluator score the results using the same rubric.
Repeat multiple rounds for each role at minimum, and report the mean, first-round pass rate, number of manual corrections, and failure types.
Compare again with candidate routes such as Sol high/medium and GPT-5.4 mini low; only codify a route when the difference is stable.
The author explicitly wrote, “These are local workflow scores, not general intelligence scores.” This is the key boundary for interpreting the result.
GPT-5.6 Terra