According to public findings from Artificial Analysis on its Intelligence Index intelligence vs. task-cost chart, GPT-5.6 Sol and Luna outperform Terra at every frontier point; consequently, Terra may lack a distinct cost/intelligence advantage, though procurement decisions should not be based solely on this post due to the lack of full underlying raw data.
Models / Interfaces: GPT-5.6 Sol, Terra, Luna, and GPT-5.5; Artificial Analysis Intelligence Index reasoning task cost vs. intelligence plot.
Input / Configuration: The original post only states that it uses the Intelligence Index chart, without disclosing the full task suite, model snapshots, reasoning effort levels, sampling parameters, pricing timestamps, or itemized sample data.
Result format: Official account textual conclusions accompanied by a chart image; no downloadable per-question itemized results were provided.
The public post does not provide directly reproducible prompts or complete configurations. The verifiable raw claim is that for reasoning tasks, the GPT-5.6 family forms the Pareto frontier relative to GPT-5.5; furthermore, for any given Terra effort level, there exists a Luna or Sol effort level that achieves higher intelligence at no additional cost, or matches its intelligence at a lower cost.
Artificial Analysis: Sol and Luna lead Terra at every point across the intelligence vs. task-cost chart.
Artificial Analysis: Luna is exceptionally cost-effective.
Rohan Paul's commentary: In paraphrasing the post, Rohan Paul stated that Terra "falls behind across the entire intelligence-cost curve," but this is a commentator's interpretation rather than an independent raw measurement.
OpenAI official page data: During the same period, OpenAI's official page listed Intelligence Index v4.1 scores as: Terra 55, Luna 51.2, Sol 58.9. Because the evaluation methodologies and measurement timestamps may differ, official aggregate scores cannot be directly conflated with the effort-level data points on Artificial Analysis's chart.
For workloads where API cost and intelligence are the primary objectives, Terra should be benchmarked against Luna and Sol within the same evaluation harness. One should not presume Terra offers optimal cost-performance simply because its naming convention places it in the mid-tier. Terra may still retain practical value in scenarios governed by specific latency profiles, context constraints, tool-calling reliability, or pre-existing team routing pipelines.
The original post did not disclose the complete chart dataset, task inventory, scoring rubrics, exact model versions, pricing baselines, or effort mappings, making it impossible to recompute individual data points.
"Leading at every point" is a curve-level aggregate conclusion and does not imply that Terra underperforms on every specific task.
Index versions and collection timelines between the official site and independent evaluators may diverge; when discrepancies arise, both original methodologies and baselines should be preserved.
Obtain the same Intelligence Index version and model snapshots, recording the collection timestamp, pricing, and reasoning effort levels.
Evaluate Terra, Luna, and Sol using an identical task suite and tooling, repeating each test at least 3 times while recording accuracy, input/output token counts, latency, and cost.
Plot the cost–intelligence curve to assess whether higher-scoring Luna or Sol configurations exist that do not exceed the cost of each corresponding Terra configuration point.
Document non-dominated points, failed tasks, and latency anomalies separately, avoiding unwarranted extrapolation of curve-level conclusions to all production workflows.
GPT-5.6 Terra