Grok 4.7 · Media / benchmark · Independent measurement
Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
Tasks suitable for a judgment: Comparing Grok 4.7 (xhigh) on the composite Intelligence Index under Artificial Analysis’s unified harness, and on the Coding Agent Index in Grok Build, its first-party coding agent; judging long-horizon knowledge work such as AA-Briefcase and GDPval-AA, plus output tokens, decode time, and hallucination rate.
Tasks unsuitable for extrapolation: Substituting the scores on this page for the self-reported benchmarks on the xAI release page; treating Grok 4.6 high and xhigh as the same tier; treating Terminal-Bench 4.0 in the Coding Agent Index as the same-named item in the Intelligence Index; reading about 7.1 minutes as end-to-end wall-clock time that includes time to first token and overhead; reading the list price as the dollar cost of this evaluation.
Applicable model versions: Grok 4.7, evaluated at xhigh. The comparisons include both Grok 4.6 (high) and Grok 4.6 (xhigh), and the two scores differ. The article does not give Grok 4.7 scores for low, medium, or high, and Grok 4.7 Fast or a fast variant does not appear.
Test environment or client: The Intelligence Index uses a unified cross-model evaluation harness; the specific software name is not stated. For the Coding Agent Index, Grok uses its first-party coding agent, Grok Build; the other models on the same chart use their own native harnesses. Visible on the chart: Claude Code, Codex, Muse Code, Opencode, Kimi Code CLI, and Antigravity SDK. API provider, region, and hardware are not stated.
Reasoning tiers and parameters: The article states that configurable reasoning effort runs from low to xhigh, and this evaluation uses xhigh. The specific meanings of temperature, top-p, random seed, maximum output, tool schema, sample size, and fallback are all unspecified.
This is an independent evaluation article from Artificial Analysis dated 2026-09-21, not an xAI release note. The body says “We evaluated the new model at xhigh reasoning effort.” The breakdown-chart title says “Intelligence evaluations measured independently by Artificial Analysis”.
The two indexes must be kept separate:
Artificial Analysis Intelligence Index v4.3. The caption states that it includes 10 items: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. The body states that this set of results uses a unified cross-model evaluation harness.
Artificial Analysis Coding Agent Index v1.5. The caption states that it includes 3 items: DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA, Higher is better. Grok’s results in this set use Grok Build. The body states that this index is separate from the Intelligence Index.
The charts also give these definitions. All of them are Artificial Analysis measurement notes, not xAI self-reports:
AA-Briefcase Elo: combines rubric pass rate, analytical quality Elo, and presentation Elo. Higher is better.
GDPval-AA v2: Elo on real work tasks, anchored to a human baseline of 1,000. Higher is better.
The three coding breakdown charts: Average pass@1. Higher is better.
Output tokens per Intelligence Index question: a weighted average. The legend separates Answer and Reasoning.
Time per question: weighted-average decode time (minutes), excludes TTFT and overhead time. Lower is better.
AA-Omniscience Index: the score ranges from -100 to 100; correct answers add points, hallucinations subtract points, and refusals are not penalized; 0 means there are as many correct answers as incorrect ones.
AA-Omniscience Accuracy: the proportion of all questions answered correctly, whether or not the model chose to answer.
AA-Omniscience Hallucination Rate: the proportion of all non-correct responses that are wrong, that is, incorrect / (incorrect + partial answers + not attempted). Lower is better.
The large breakdown chart at the end plots AA-Briefcase and GDPval-AA v2 as (Elo-500)/2000, so the percentages in those two cells are not raw Elo. Terminal-Bench 4.0 and GDP.pdf are marked NEW on that chart. The GDP.pdf subtitle is “Professional document reasoning, All-pass”.
The context window and list prices under “Other model details” are model specifications listed in the article. They are not scores on the Intelligence Index or the Coding Agent Index.
Below, Artificial Analysis measurements are listed first, then the model specifications in the article.
Artificial Analysis measurements
Intelligence Index: Grok 4.7 is 46. The body says this is 2 points higher than Grok 4.6. The index chart labels Grok 4.7 as (xhigh), 46, and the Grok 4.6 beside it as (xhigh), 44.
Coding agent: Grok 4.7 (xhigh) + Grok Build is 56, 9 points higher than Grok 4.6 (xhigh) + Grok Build at 47. The body says it ranks 4th in the native-harness comparison, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. The headline says the Coding Agent Index surpasses GPT-5.6 Sol.
AA-Briefcase: Grok 4.7 is 1657 Elo, 111 higher than Grok 4.6 (high). Analytical quality is 1994 Elo and presentation quality is 1499 Elo; the corresponding figures for Grok 4.6 (high) are 1690 and 1519. Presentation quality is lower than Grok 4.6 (high).
GDPval-AA: Grok 4.7 is 1695 Elo and Grok 4.6 (high) is 1605 Elo. The body says this is 90 higher.
Inside the Intelligence Index, under the unified harness, relative to Grok 4.6 (high): Terminal-Bench 4.0 is 4.5 percentage points higher, GDP.pdf is 3.0 percentage points higher, AA-LCR is 3.7 percentage points lower, and AutomationBench-AA is 1.1 percentage points lower. The body does not give absolute scores for these four items.
Grok Build component scores, which are not the same set of numbers as the unified harness above: DeepSWE v1.1 goes from 65% to 73%, Terminal-Bench 4.0 from 18% to 33%, and SWE-Atlas-QnA from 58% to 63%. The comparison model is Grok 4.6 (xhigh).
Output tokens: Grok 4.7 (xhigh) uses about 81k per Intelligence Index question. The body first compares this with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max), and says the increases are 125% and 196%, respectively. A later passage also says Grok 4.6 (xhigh) is 38k, Muse Spark 1.3 (max) is 60k, and GPT-6 Astra (max) is still 27k.
Speed and time: on long prompts, answer output speed is about 188 tokens/second. Each Intelligence Index question takes about 7.1 minutes; the time chart defines this metric as decode time excluding TTFT and overhead.
AA-Omniscience: the hallucination rate for Grok 4.7 (xhigh) is 29%, and for Grok 4.6 (high) it is 34%. Accuracy is 47% versus 48%. The index goes from 30 to 32.
Model specifications listed in the article, not scores from this evaluation
The context window is 500k tokens, the same as Grok 4.6.
The price is $2 input and $6 output per million tokens, and a cache hit is $0.50 per million tokens, the same as Grok 4.6.
Reasoning effort can be set from low to xhigh.
The headline also says Grok 4.7 puts SpaceXAI into the “top 4 AI labs”. The article does not name those four labs. On the index chart, Grok 4.7 (xhigh) at 46 is below Claude Fable 5.1, GPT-6 Astra, Claude Opus 5, Claude Fable 5, Muse Spark 1.3, and GPT-5.6 Sol.
The collection date is 2026-09-22. The Elo charts have error bars, but the interval endpoints are not printed, so the tables below record only the point estimates on the bars. Reasoning tiers that are not printed are not filled in.
The 10 items listed in the caption are given above. Bar order matches the chart from left to right.
| Model and tier | Index |
|---|---|
| Claude Fable 5.1 (max with fallback) | 53 |
| GPT-6 Astra (max) | 53 |
| Claude Opus 5 (max) | 51 |
| Claude Fable 5 (with fallback) | 50 |
| Muse Spark 1.3 (max) | 48 |
| GPT-5.6 Sol (max) | 47 |
| Grok 4.7 (xhigh) | 46 |
| Qwen3.8 Max (0902) | 45 |
| GLM-5.3 (max) | 45 |
| Grok 4.6 (xhigh) | 44 |
| Step 5 Preview | 44 |
| Kimi K3 (max) | 44 |
| GPT-5.6 Terra (max) | 42 |
| GLM-5.3-Flash | 42 |
| Gemini 3.8 Flash (high) | 41 |
| DeepSeek V4.1 Flash (max) | 39 |
| GPT-5.6 Luna (max) | 37 |
| DeepSeek V4 Pro 0813 (max) | 36 |
| Qwen3.8 27B (xhigh) | 34 |
| K2-Horizon 375B A23B | 31 |
| MiniMax-M3 | 29 |
| Inkling | 25 |
| Nemotron 3 Ultra | 23 |
| Gemini 3.5 Flash-Lite | 22 |
| Muse Glimmer (high) | 17 |
| Mistral Medium 3.5 | 14 |
| gpt-oss-120b (high) | 12 |
Harness and model are split into rows following the labels on the chart. Grok uses Grok Build.
| Harness | Model and tier | Index |
|---|---|---|
| Claude Code | Fable 5.1 (max) (with fallback) | 62 |
| Codex | GPT-6 Astra (max) | 62 |
| Claude Code | Opus 5 (max) | 60 |
| Grok Build | Grok 4.7 (xhigh) | 56 |
| Codex | GPT-5.6 Sol (max) | 55 |
| Muse Code | Muse Spark 1.3 (max) | 54 |
| Opencode | GLM-5.3 | 54 |
| Kimi Code CLI | Kimi K3 | 52 |
| Grok Build | Grok 4.6 (xhigh) | 47 |
| Claude Code | Qwen3.8 Max | 43 |
| Codex | DeepSeek V4 Pro 0813 (max) | 43 |
| Antigravity SDK | Gemini 3.8 Flash (high) | 42 |
The three Grok Build pass@1 scores. The body gives Grok 4.7 (xhigh) relative to Grok 4.6 (xhigh):
| Benchmark | Grok 4.6 (xhigh) | Grok 4.7 (xhigh) |
|---|---|---|
| DeepSWE v1.1 | 65% | 73% |
| Terminal-Bench 4.0 | 18% | 33% |
| SWE-Atlas-QnA | 58% | 63% |
On the AA-Briefcase chart, Grok 4.7 is (xhigh). The GDPval chart likewise labels Grok 4.7 as (xhigh). The human baseline is 1,000.
| Rank | AA-Briefcase model and tier | Elo | GDPval-AA v2 model and tier | Elo |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 (max with fallback) | 1678 | Claude Fable 5.1 (max with fallback) | 1735 |
| 2 | Claude Opus 5 (max) | 1673 | Claude Opus 5 (max) | 1708 |
| 3 | Grok 4.7 (xhigh) | 1657 | Grok 4.7 (xhigh) | 1695 |
| 4 | Qwen3.8 Max (0902) | 1640 | Muse Spark 1.3 (max) | 1674 |
| 5 | Muse Spark 1.3 (max) | 1597 | Qwen3.8 Max (0902) | 1668 |
| 6 | GPT-6 Astra (max) | 1569 | GLM-5.3 (max) | 1646 |
| 7 | Grok 4.6 (high) | 1546 | GLM-5.3-Flash | 1641 |
| 8 | Claude Fable 5 (with fallback) | 1543 | Grok 4.6 (high) | 1605 |
| 9 | GLM-5.3 (max) | 1525 | DeepSeek V4.1 Flash (max) | 1600 |
| 10 | Kimi K3 (max) | 1510 | Claude Fable 5 (with fallback) | 1595 |
| 11 | GPT-5.6 Sol (max) | 1487 | GPT-5.6 Sol (max) | 1588 |
| 12 | GLM-5.3-Flash | 1459 | Step 5 Preview | 1566 |
| 13 | Step 5 Preview | 1432 | GPT-6 Astra (max) | 1542 |
| 14 | DeepSeek V4.1 Flash (max) | 1432 | Kimi K3 (max) | 1524 |
| 15 | Qwen3.8 27B (xhigh) | 1403 | GPT-5.6 Luna (max) | 1443 |
| 16 | GPT-5.6 Luna (max) | 1345 | DeepSeek V4 Pro 0813 (max) | 1441 |
| 17 | GPT-5.6 Terra (max) | 1336 | GPT-5.6 Terra (max) | 1432 |
| 18 | K2-Horizon 375B A23B | 1303 | Gemini 3.8 Flash (high) | 1412 |
| 19 | DeepSeek V4 Pro 0813 (max) | 1261 | Qwen3.8 27B (xhigh) | 1409 |
| 20 | Gemini 3.8 Flash (high) | 1202 | K2-Horizon 375B A23B | 1349 |
| 21 | MiniMax-M3 | 1092 | MiniMax-M3 | 1230 |
| 22 | Nemotron 3 Ultra | 873 | Inkling | 1064 |
| 23 | Inkling | 829 | Nemotron 3 Ultra | 1000 |
| 24 | Gemini 3.5 Flash-Lite | 642 | Gemini 3.5 Flash-Lite | 970 |
| 25 | Mistral Medium 3.5 | 515 | Muse Glimmer (high) | 774 |
| 26 | Muse Glimmer (high) | 469 | Mistral Medium 3.5 | 747 |
| 27 | gpt-oss-120b (high) | 0 | gpt-oss-120b (high) | 596 |
The body also gives AA-Briefcase quality scores, and they cover only this pair:
| Metric | Grok 4.7 | Grok 4.6 (high) |
|---|---|---|
| Overall Elo | 1657 | 1546 |
| Analytical quality Elo | 1994 | 1690 |
| Presentation quality Elo | 1499 | 1519 |
The body does not separately repeat the reasoning tier for Grok 4.7 in this quality-score table. The overall Elo of 1657 in the same passage matches the 1657 for (xhigh) in the table above.
Output tokens per Intelligence Index question. The body uses “approximately”. The chart also prints a total and splits Answer and Reasoning into two segments. Adding 11k and 17k does not equal the 27k at the top of the bar. The table keeps the numbers printed on the chart and does not rewrite them into values that add up.
| Model and tier | Approximate figure in the body | Total on the chart | Answer on the chart | Reasoning on the chart |
|---|---|---|---|---|
| GPT-6 Astra (max) | 27k | 27k | 11k | 17k |
| Grok 4.6 (high) | 36k | 36k | 17k | 19k |
| Grok 4.6 (xhigh) | 38k | 38k | 18k | 20k |
| Muse Spark 1.3 (max) | 60k | 60k | 26k | 34k |
| Grok 4.7 (xhigh) | 81k | 81k | 22k | 59k |
The body says 81k is 125% more than 36k and 196% more than 27k.
| Metric | Value | Definition |
|---|---|---|
| Answer output speed | about 188 tokens/second | Body: “for long prompts”. Prompt length is not stated |
| Time per Intelligence Index question | about 7.1 minutes | Time chart: weighted-average decode time, excluding TTFT and overhead |
| Metric | Grok 4.7 (xhigh) | Grok 4.6 (high) |
|---|---|---|
| Index | 32 | 30 |
| Accuracy | 47% | 48% |
| Hallucination Rate | 29% | 34% |
| Item | Wording in the article |
|---|---|
| Context window | 500k tokens, the same as Grok 4.6 |
| Input / output price | $2 / $6 per million tokens |
| Cache hit | $0.50 per million tokens |
| Reasoning-effort range | low to xhigh |
| Tier used in this evaluation | xhigh |
| Dollar cost per evaluation question | not stated |
In this article, Grok 4.7 (xhigh)’s advantages concentrate on two kinds of tasks: long-horizon knowledge work under the unified harness (AA-Briefcase 1657, GDPval-AA 1695), and the coding-agent index on Grok Build (56). Analytical quality Elo is clearly higher than Grok 4.6 (high), and presentation quality Elo is lower. Under the unified harness, Terminal-Bench 4.0 and GDP.pdf rise slightly, and AA-LCR and AutomationBench-AA fall slightly. The body does not give absolute scores for these changes.
Token use is part of the limitation. About 81k output tokens is higher than the figures the body names: 36k for Grok 4.6 (high), 38k for Grok 4.6 (xhigh), 60k for Muse Spark 1.3 (max), and 27k for GPT-6 Astra (max). Of the 81k shown for Grok 4.7 on the chart, the Reasoning segment is 59k and the Answer segment is 22k.
The following cannot be folded into the measurements above:
Self-reported scores on the xAI release page, such as CursorBench, DeepSWE, and EEBench. This note did not open that page or cite its numbers. If DeepSWE or Terminal-Bench is cited, only this article’s Artificial Analysis / Grok Build definitions may be used.
The model index page. The article links to that page; this note does not use it to fill in scores, provider, or price.
high and xhigh. The +2 on the Intelligence Index chart corresponds to 44 for Grok 4.6 (xhigh). The AA-Briefcase, GDPval, and Omniscience comparisons are against Grok 4.6 (high). The coding-agent comparison is against Grok 4.6 (xhigh).
The two Terminal-Bench 4.0 results. The unified harness gives only +4.5 percentage points relative to Grok 4.6 (high). Grok Build gives 18% to 33%.
7.1 minutes. This is decode time excluding TTFT and overhead.
$2 / $6 and cache $0.50. These are list prices. The article does not give the dollar cost of running the Intelligence Index or the coding index.
“top 4 AI labs”. This is headline wording. The article does not provide a lab ranking table.
Sample size, confidence-interval figures, temperature, seed, provider, API model ID, and the meaning of “with fallback” are all unspecified.
Open https://artificialanalysis.ai/articles/benchmarking-grok-4-7 and check the date September 21, 2026 and the index of 46 in the headline.
Read the body first, and keep xhigh, Grok Build, and the unified harness separate. Do not use the model index page to fill in fields this article does not have.
Index scores, Elo, tokens, and coding-agent scores are in the images inside the article. Read the numbers at the tops of the bars and the angled model names, and keep the reasoning tiers printed on the chart. For the Elo charts, restate only the point estimates.
On the large breakdown chart, the AA-Briefcase and GDPval axes are (Elo-500)/2000. Do not treat those percentages as raw Elo, and do not treat them as the 1657 or 1695 in the body.
The article does not publish a question list, scoring script, random seed, or request parameters, so a third party cannot rerun it question by question from this article. What can be checked is the index versions, harness names, tiers, Elo scores, token counts, times, and the comparison-model set on this 2026-09-22 version of the page.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Artificial Analysis · Artificial Analysis (the author type in the page JSON-LD is Organization) · Original publication date 2026-09-21 · Site edit date 2026-09-22
Open original sourceGrok 4.7
Download the Tabbit client to check model access