Grok 4.6 · Media / benchmark · Editorial analysis
Article conclusion KIE summarizes Grok 4.6's Intelligence Index as 61, tied with GPT-5.6 Sol; its price is $2 for input and $6 for output per 1M tokens. The article argues that its main selling point is its price-to-intelligence ratio, rather than leading on e。
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
KIE summarizes Grok 4.6's Intelligence Index as 61, tied with GPT-5.6 Sol; its price is $2 for input and $6 for output per 1M tokens. The article argues that its main selling point is its price-to-intelligence ratio, rather than leading on every individual evaluation.
Grok 4.6's strengths are concentrated in knowledge work and legal reasoning: GDPVal-AA 1753, AA-Briefcase 1577, and Harvey LAB 15.8%.
Its Terminal-Bench v3.0 score is 26%, below GPT-5.6 Sol's 34.6%, making it one of the clearest weaknesses in the release table.
Its DeepSWE v1.1 score is 65.9%, a substantial increase from Grok 4.5's 54%, but still below Sol's 73%.
APEX-Agents rose from 47.1% to 57.5%, indicating clear progress on long-horizon Agent tasks.
The article emphasizes that, as of 2026-08-13, the public data still contained no complete independent third-party reproduction. The competing results in the table come from each model's publicly reported scores and cannot replace measurements under a unified harness.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
KIE.ai Blog · Sofia Marenco (Model Evaluation Lead) · Original publication date 2026-08-13 · Site edit date 2026-09-20
Open original sourceGrok 4.6
Download the Tabbit client to check model access