Grok 4.6 is positioned for long-running agent tasks, coding, knowledge work, and interactive/visual projects. The company says it ties with GPT-5.6 Sol on the Artificial Analysis Intelligence Index, with an overall score of 61; however, the official table also shows it performing better on knowledge-work evaluations while trailing GPT-5.6 Sol Max on software-engineering evaluations such as DeepSWE and Terminal-Bench.
| Evaluation | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (Extended) | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | — | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
Maintain context over long-running tasks and continue making progress across research, repository operations, and application building.
Perform strongly on knowledge-work and legal-work evaluations such as GDPVal-AA v2, AA-Briefcase, and Harvey LAB.
Training covers general coding, knowledge work, kernel optimization, web development, CAD, and other agent environments.
The company says more self-testing and verification behavior appears in long trajectories.
Context window: 500,000 tokens
Input price: $2 / 1M tokens
Output price: $6 / 1M tokens
Available through: Cursor, Grok Build, API, OpenRouter, Vercel, Cloudflare, and others
The company notes that competitor scores come from system cards or public leaderboards published by their respective developers, rather than from four models rerun in the same experimental environment. The table is therefore suitable for directional judgment, but should not be treated as a strict head-to-head experimental conclusion. In particular, the Terminal-Bench version, reasoning level, and execution harness all affect the results.
Grok 4.6