Emergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61, tied with GPT-5.6 Sol; but on DeepSWE v1.1 and Terminal-Bench v3.0, it trails GPT-5.6 Sol by 7.1 and 8.6 percentage points, respectively.
| Evaluation | Grok 4.6 | Grok 4.5 | GPT-5.6 Sol | Fable 5 |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPval-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB | 15.8% | 12.9% | 2.5% | 11.3% |
Grok 4.6 scored highest in the table on GDPval-AA, AA-Briefcase, and the Harvey legal-work evaluation, suggesting that it is better suited to research, analysis, documentation, and professional knowledge work.
Compared with Grok 4.5, it improved on every publicly reported row, with especially notable gains on Agent tasks.
Artificial Analysis additionally measured a 65.7% non-hallucination rate. This metric matters for customer service, knowledge bases, and user-facing Agents, but it should not be treated as equivalent to "factual accuracy."
The article notes that xAI's table juxtaposes "self-reported or publicly available best scores" for each model; it is not a strictly controlled head-to-head retest. The difference between Terminal-Bench v2.1 and v3.0 scores also shows that benchmark citations must specify the version, reasoning tier, and test source.
Grok 4.6