The AI IQ model page gives GPT-6 Sol an estimated overall IQ of 136, but only academic reasoning, coding reasoning, and reliability have direct benchmark results among the six dimensions. The remaining dimensions and some missing benchmarks are estimated or imputed by the scoring process. The score of 136 should therefore be understood as a model estimate with incomplete coverage, not as a directly measured human-IQ equivalent.
| Dimension | IQ shown on page | Coverage | Benchmark results listed on page |
|---|---|---|---|
| Abstract reasoning | 131 | 0/3 | None |
| Mathematical reasoning | 139 | 0/5 | None |
| Academic reasoning | 141 | 4/6 | CritPt 30.8571; Humanity’s Last Exam 47.9147; MMMU-Pro 83.2948; SciCode 57.6389 |
| Coding reasoning | 143 | 1/6 | Terminal-Bench 4.0 43.9394 |
| Computer use | 137 | 0/6 | None |
| Reliability | 124 | 2/7 | AA Long Context Reasoning v1.1 83.6667; AA Omniscience 27.1167 |
The model page also lists an overall IQ of 136, an IQ rank of #6, and an effective cost of $9.0775. The individual benchmark table shows values without specifying units; the values shown on the site are preserved here without converting them to percentages or other units. Effective cost is not a model capability score.
The methodology page explains:
Overall IQ is the average of six equally weighted dimension scores. Raw benchmark scores are first mapped to IQ values using each benchmark's calibration ladder. One benchmark with a source can produce a dimension estimate; broader coverage increases confidence in the estimate.
Missing benchmarks and entirely missing dimensions are conservatively imputed in the scoring process. The coverage counts shown on the page distinguish direct benchmark coverage from estimated portions; a full dimension score does not mean that the dimension was measured.
The site publishes an overall IQ only when at least two dimensions have supporting sources. The GPT-6 Sol page has benchmark results in three dimensions, while the other three have 0 coverage. The overall score of 136 therefore includes estimates for missing coverage.
The methodology page says that both dimension and overall scores are estimates; even with complete coverage, they are not a human psychometric test or an IQ study conducted on the model. Score mapping depends on the benchmark calibration ladders defined by the site, and newly added benchmarks may use provisional mappings.
These data are suitable as a third-party aggregated reference within AI IQ's own scoring system. They do not support a claim that GPT-6 Sol's “true IQ” is 136, nor should the dimension scores be treated as direct measurements.
The model page aggregates benchmark results but does not provide, for each evaluation, the evaluating organization, the specific model version and reasoning level, prompts, tools or harness, sample size, or confidence interval. The page alone is insufficient to independently rerun or verify every raw score.
Coverage is zero in several dimensions; coding reasoning has only 1/6 coverage, and reliability has 2/7. Coverage is sparse, so the overall score is affected by estimates and missing-value handling.
The individual scores in the model-page table have no units specified. If cited, they should be kept as displayed and this limitation should be stated.
Rankings can change as the site adds models and benchmarks or updates its calibration. This is a ranking within AI IQ, not a standardized cross-site ranking.
The primary evidence is the GPT-6 Sol model profile, which lists the overall IQ, six dimension IQ scores, dimension coverage counts, and seven benchmark scores. The AI IQ methodology page provides additional details on the scoring rules. This record summarizes only information visible on those two pages; it does not present estimates as direct measurements or infer results for unlisted benchmarks.
GPT-6 Sol