Grok 4.6 · Media / benchmark · Independent measurement
Emergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61, tied with GPT-5.6 Sol; but on DeepSWE 。
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
Emergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61, tied with GPT-5.6 Sol; but on DeepSWE v1.1 and Terminal-Bench v3.0, it trails GPT-5.6 Sol by 7.1 and 8.6 percentage points, respectively.
| Evaluation | Grok 4.6 | Grok 4.5 | GPT-5.6 Sol | Fable 5 |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPval-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB | 15.8% | 12.9% | 2.5% | 11.3% |
Grok 4.6 scored highest in the table on GDPval-AA, AA-Briefcase, and the Harvey legal-work evaluation, suggesting that it is better suited to research, analysis, documentation, and professional knowledge work.
Compared with Grok 4.5, it improved on every publicly reported row, with especially notable gains on Agent tasks.
Artificial Analysis additionally measured a 65.7% non-hallucination rate. This metric matters for customer service, knowledge bases, and user-facing Agents, but it should not be treated as equivalent to "factual accuracy."
The article notes that xAI's table juxtaposes "self-reported or publicly available best scores" for each model; it is not a strictly controlled head-to-head retest. The difference between Terminal-Bench v2.1 and v3.0 scores also shows that benchmark citations must specify the version, reasoning tier, and test source.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Emergent Learn · Anupam; reviewed by Anmol · Original publication date 2026-08-13 · Site edit date 2026-09-20
Open original sourceGrok 4.6
Download the Tabbit client to check model access