Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaGPT-6 Sol

GPT-6 Sol on AI IQ: Model Profile and Benchmark Coverage

Original source

AI IQ

AuthorLiberated Software LLC (site attribution)

Source date2026-09-22

Tabbit curation2026-09-22

Read original

One-sentence takeaway

The AI IQ model page gives GPT-6 Sol an estimated overall IQ of 136, but only academic reasoning, coding reasoning, and reliability have direct benchmark results among the six dimensions. The remaining dimensions and some missing benchmarks are estimated or imputed by the scoring process. The score of 136 should therefore be understood as a model estimate with incomplete coverage, not as a directly measured human-IQ equivalent.

Model-page data

Six dimension scores and benchmark coverage

DimensionIQ shown on pageCoverageBenchmark results listed on page
Abstract reasoning1310/3None
Mathematical reasoning1390/5None
Academic reasoning1414/6CritPt 30.8571; Humanity’s Last Exam 47.9147; MMMU-Pro 83.2948; SciCode 57.6389
Coding reasoning1431/6Terminal-Bench 4.0 43.9394
Computer use1370/6None
Reliability1242/7AA Long Context Reasoning v1.1 83.6667; AA Omniscience 27.1167

The model page also lists an overall IQ of 136, an IQ rank of #6, and an effective cost of $9.0775. The individual benchmark table shows values without specifying units; the values shown on the site are preserved here without converting them to percentages or other units. Effective cost is not a model capability score.

IQ estimation and limits of missing coverage

The methodology page explains:

  • Overall IQ is the average of six equally weighted dimension scores. Raw benchmark scores are first mapped to IQ values using each benchmark's calibration ladder. One benchmark with a source can produce a dimension estimate; broader coverage increases confidence in the estimate.

  • Missing benchmarks and entirely missing dimensions are conservatively imputed in the scoring process. The coverage counts shown on the page distinguish direct benchmark coverage from estimated portions; a full dimension score does not mean that the dimension was measured.

  • The site publishes an overall IQ only when at least two dimensions have supporting sources. The GPT-6 Sol page has benchmark results in three dimensions, while the other three have 0 coverage. The overall score of 136 therefore includes estimates for missing coverage.

  • The methodology page says that both dimension and overall scores are estimates; even with complete coverage, they are not a human psychometric test or an IQ study conducted on the model. Score mapping depends on the benchmark calibration ladders defined by the site, and newly added benchmarks may use provisional mappings.

These data are suitable as a third-party aggregated reference within AI IQ's own scoring system. They do not support a claim that GPT-6 Sol's “true IQ” is 136, nor should the dimension scores be treated as direct measurements.

Limitations

  • The model page aggregates benchmark results but does not provide, for each evaluation, the evaluating organization, the specific model version and reasoning level, prompts, tools or harness, sample size, or confidence interval. The page alone is insufficient to independently rerun or verify every raw score.

  • Coverage is zero in several dimensions; coding reasoning has only 1/6 coverage, and reliability has 2/7. Coverage is sparse, so the overall score is affected by estimates and missing-value handling.

  • The individual scores in the model-page table have no units specified. If cited, they should be kept as displayed and this limitation should be stated.

  • Rankings can change as the site adds models and benchmarks or updates its calibration. This is a ranking within AI IQ, not a standardized cross-site ranking.

Raw evidence and data

The primary evidence is the GPT-6 Sol model profile, which lists the overall IQ, six dimension IQ scores, dimension coverage counts, and seven benchmark scores. The AI IQ methodology page provides additional details on the scoring rules. This record summarizes only information visible on those two pages; it does not present estimates as direct measurements or infer results for unlisted benchmarks.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-6 Sol

Use and compare models in Tabbit

GPT-6 Sol

Related reviews

OfficialOpenAI2026-09-22

GPT-6 Sol: Official Benchmarks and Evaluation Boundaries

MediaArtificial Analysis2026-09-22

GPT-6 Sol: Artificial Analysis on Cost Efficiency and Hallucination Measurement

CommunityKillSwitch-Bench

GPT-6 Sol: KillSwitch-Bench Adversarial Esoteric-Language Coding Agent Benchmark

MediaArtificial Analysis2026-09

GPT-6 Sol: Artificial Analysis Comparison Across Six Configurations

GPT-6 Sol

Related prompts

OfficialOpenAI Developers

GPT-6 Sol Official API Model Configuration

OfficialOpenAI Developers

OpenAI GPT-6 Family Prompting Guide

OfficialOpenAI official release notes2026-09-22

GPT-6 Prompt Caching Optimization Workflow

OfficialOpenAI Developers

OpenAI's Official GPT-6 Async Tool-Calling Workflow