Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Grok 4.7 · Media / benchmark · Independent measurement

Artificial Analysis: Grok 4.7 Intelligence Index and Coding Agent

Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.

Media / benchmarkIndependent measurementEdited 2026-09-22

Test conditions

Source-specific observation
Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
Published conditions
Substituting the scores on this page for the self-reported benchmarks on the xAI release page; treating Grok 4.6 high and xhigh as the same tier; treating Terminal-Bench 4.0 in the Coding Agent Index as the same-named item in the Intelligence Index; reading about 7.1 minutes as end-to-end wall-clock time that includes。

Key data and applicable tasks

One-sentence takeaway

Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.

Use cases

  • Tasks suitable for a judgment: Comparing Grok 4.7 (xhigh) on the composite Intelligence Index under Artificial Analysis’s unified harness, and on the Coding Agent Index in Grok Build, its first-party coding agent; judging long-horizon knowledge work such as AA-Briefcase and GDPval-AA, plus output tokens, decode time, and hallucination rate.

  • Tasks unsuitable for extrapolation: Substituting the scores on this page for the self-reported benchmarks on the xAI release page; treating Grok 4.6 high and xhigh as the same tier; treating Terminal-Bench 4.0 in the Coding Agent Index as the same-named item in the Intelligence Index; reading about 7.1 minutes as end-to-end wall-clock time that includes time to first token and overhead; reading the list price as the dollar cost of this evaluation.

  • Applicable model versions: Grok 4.7, evaluated at xhigh. The comparisons include both Grok 4.6 (high) and Grok 4.6 (xhigh), and the two scores differ. The article does not give Grok 4.7 scores for low, medium, or high, and Grok 4.7 Fast or a fast variant does not appear.

  • Test environment or client: The Intelligence Index uses a unified cross-model evaluation harness; the specific software name is not stated. For the Coding Agent Index, Grok uses its first-party coding agent, Grok Build; the other models on the same chart use their own native harnesses. Visible on the chart: Claude Code, Codex, Muse Code, Opencode, Kimi Code CLI, and Antigravity SDK. API provider, region, and hardware are not stated.

  • Reasoning tiers and parameters: The article states that configurable reasoning effort runs from low to xhigh, and this evaluation uses xhigh. The specific meanings of temperature, top-p, random seed, maximum output, tool schema, sample size, and fallback are all unspecified.

Evaluation method

This is an independent evaluation article from Artificial Analysis dated 2026-09-21, not an xAI release note. The body says “We evaluated the new model at xhigh reasoning effort.” The breakdown-chart title says “Intelligence evaluations measured independently by Artificial Analysis”.

The two indexes must be kept separate:

  • Artificial Analysis Intelligence Index v4.3. The caption states that it includes 10 items: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. The body states that this set of results uses a unified cross-model evaluation harness.

  • Artificial Analysis Coding Agent Index v1.5. The caption states that it includes 3 items: DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA, Higher is better. Grok’s results in this set use Grok Build. The body states that this index is separate from the Intelligence Index.

The charts also give these definitions. All of them are Artificial Analysis measurement notes, not xAI self-reports:

  • AA-Briefcase Elo: combines rubric pass rate, analytical quality Elo, and presentation Elo. Higher is better.

  • GDPval-AA v2: Elo on real work tasks, anchored to a human baseline of 1,000. Higher is better.

  • The three coding breakdown charts: Average pass@1. Higher is better.

  • Output tokens per Intelligence Index question: a weighted average. The legend separates Answer and Reasoning.

  • Time per question: weighted-average decode time (minutes), excludes TTFT and overhead time. Lower is better.

  • AA-Omniscience Index: the score ranges from -100 to 100; correct answers add points, hallucinations subtract points, and refusals are not penalized; 0 means there are as many correct answers as incorrect ones.

  • AA-Omniscience Accuracy: the proportion of all questions answered correctly, whether or not the model chose to answer.

  • AA-Omniscience Hallucination Rate: the proportion of all non-correct responses that are wrong, that is, incorrect / (incorrect + partial answers + not attempted). Lower is better.

The large breakdown chart at the end plots AA-Briefcase and GDPval-AA v2 as (Elo-500)/2000, so the percentages in those two cells are not raw Elo. Terminal-Bench 4.0 and GDP.pdf are marked NEW on that chart. The GDP.pdf subtitle is “Professional document reasoning, All-pass”.

The context window and list prices under “Other model details” are model specifications listed in the article. They are not scores on the Intelligence Index or the Coding Agent Index.

Key results

Below, Artificial Analysis measurements are listed first, then the model specifications in the article.

Artificial Analysis measurements

  • Intelligence Index: Grok 4.7 is 46. The body says this is 2 points higher than Grok 4.6. The index chart labels Grok 4.7 as (xhigh), 46, and the Grok 4.6 beside it as (xhigh), 44.

  • Coding agent: Grok 4.7 (xhigh) + Grok Build is 56, 9 points higher than Grok 4.6 (xhigh) + Grok Build at 47. The body says it ranks 4th in the native-harness comparison, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. The headline says the Coding Agent Index surpasses GPT-5.6 Sol.

  • AA-Briefcase: Grok 4.7 is 1657 Elo, 111 higher than Grok 4.6 (high). Analytical quality is 1994 Elo and presentation quality is 1499 Elo; the corresponding figures for Grok 4.6 (high) are 1690 and 1519. Presentation quality is lower than Grok 4.6 (high).

  • GDPval-AA: Grok 4.7 is 1695 Elo and Grok 4.6 (high) is 1605 Elo. The body says this is 90 higher.

  • Inside the Intelligence Index, under the unified harness, relative to Grok 4.6 (high): Terminal-Bench 4.0 is 4.5 percentage points higher, GDP.pdf is 3.0 percentage points higher, AA-LCR is 3.7 percentage points lower, and AutomationBench-AA is 1.1 percentage points lower. The body does not give absolute scores for these four items.

  • Grok Build component scores, which are not the same set of numbers as the unified harness above: DeepSWE v1.1 goes from 65% to 73%, Terminal-Bench 4.0 from 18% to 33%, and SWE-Atlas-QnA from 58% to 63%. The comparison model is Grok 4.6 (xhigh).

  • Output tokens: Grok 4.7 (xhigh) uses about 81k per Intelligence Index question. The body first compares this with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max), and says the increases are 125% and 196%, respectively. A later passage also says Grok 4.6 (xhigh) is 38k, Muse Spark 1.3 (max) is 60k, and GPT-6 Astra (max) is still 27k.

  • Speed and time: on long prompts, answer output speed is about 188 tokens/second. Each Intelligence Index question takes about 7.1 minutes; the time chart defines this metric as decode time excluding TTFT and overhead.

  • AA-Omniscience: the hallucination rate for Grok 4.7 (xhigh) is 29%, and for Grok 4.6 (high) it is 34%. Accuracy is 47% versus 48%. The index goes from 30 to 32.

Model specifications listed in the article, not scores from this evaluation

  • The context window is 500k tokens, the same as Grok 4.6.

  • The price is $2 input and $6 output per million tokens, and a cache hit is $0.50 per million tokens, the same as Grok 4.6.

  • Reasoning effort can be set from low to xhigh.

The headline also says Grok 4.7 puts SpaceXAI into the “top 4 AI labs”. The article does not name those four labs. On the index chart, Grok 4.7 (xhigh) at 46 is below Claude Fable 5.1, GPT-6 Astra, Claude Opus 5, Claude Fable 5, Muse Spark 1.3, and GPT-5.6 Sol.

Raw data

The collection date is 2026-09-22. The Elo charts have error bars, but the interval endpoints are not printed, so the tables below record only the point estimates on the bars. Reasoning tiers that are not printed are not filled in.

1. Intelligence Index v4.3

The 10 items listed in the caption are given above. Bar order matches the chart from left to right.

Model and tierIndex
Claude Fable 5.1 (max with fallback)53
GPT-6 Astra (max)53
Claude Opus 5 (max)51
Claude Fable 5 (with fallback)50
Muse Spark 1.3 (max)48
GPT-5.6 Sol (max)47
Grok 4.7 (xhigh)46
Qwen3.8 Max (0902)45
GLM-5.3 (max)45
Grok 4.6 (xhigh)44
Step 5 Preview44
Kimi K3 (max)44
GPT-5.6 Terra (max)42
GLM-5.3-Flash42
Gemini 3.8 Flash (high)41
DeepSeek V4.1 Flash (max)39
GPT-5.6 Luna (max)37
DeepSeek V4 Pro 0813 (max)36
Qwen3.8 27B (xhigh)34
K2-Horizon 375B A23B31
MiniMax-M329
Inkling25
Nemotron 3 Ultra23
Gemini 3.5 Flash-Lite22
Muse Glimmer (high)17
Mistral Medium 3.514
gpt-oss-120b (high)12

2. Coding Agent Index v1.5

Harness and model are split into rows following the labels on the chart. Grok uses Grok Build.

HarnessModel and tierIndex
Claude CodeFable 5.1 (max) (with fallback)62
CodexGPT-6 Astra (max)62
Claude CodeOpus 5 (max)60
Grok BuildGrok 4.7 (xhigh)56
CodexGPT-5.6 Sol (max)55
Muse CodeMuse Spark 1.3 (max)54
OpencodeGLM-5.354
Kimi Code CLIKimi K352
Grok BuildGrok 4.6 (xhigh)47
Claude CodeQwen3.8 Max43
CodexDeepSeek V4 Pro 0813 (max)43
Antigravity SDKGemini 3.8 Flash (high)42

The three Grok Build pass@1 scores. The body gives Grok 4.7 (xhigh) relative to Grok 4.6 (xhigh):

BenchmarkGrok 4.6 (xhigh)Grok 4.7 (xhigh)
DeepSWE v1.165%73%
Terminal-Bench 4.018%33%
SWE-Atlas-QnA58%63%

3. AA-Briefcase Elo and GDPval-AA v2

On the AA-Briefcase chart, Grok 4.7 is (xhigh). The GDPval chart likewise labels Grok 4.7 as (xhigh). The human baseline is 1,000.

RankAA-Briefcase model and tierEloGDPval-AA v2 model and tierElo
1Claude Fable 5.1 (max with fallback)1678Claude Fable 5.1 (max with fallback)1735
2Claude Opus 5 (max)1673Claude Opus 5 (max)1708
3Grok 4.7 (xhigh)1657Grok 4.7 (xhigh)1695
4Qwen3.8 Max (0902)1640Muse Spark 1.3 (max)1674
5Muse Spark 1.3 (max)1597Qwen3.8 Max (0902)1668
6GPT-6 Astra (max)1569GLM-5.3 (max)1646
7Grok 4.6 (high)1546GLM-5.3-Flash1641
8Claude Fable 5 (with fallback)1543Grok 4.6 (high)1605
9GLM-5.3 (max)1525DeepSeek V4.1 Flash (max)1600
10Kimi K3 (max)1510Claude Fable 5 (with fallback)1595
11GPT-5.6 Sol (max)1487GPT-5.6 Sol (max)1588
12GLM-5.3-Flash1459Step 5 Preview1566
13Step 5 Preview1432GPT-6 Astra (max)1542
14DeepSeek V4.1 Flash (max)1432Kimi K3 (max)1524
15Qwen3.8 27B (xhigh)1403GPT-5.6 Luna (max)1443
16GPT-5.6 Luna (max)1345DeepSeek V4 Pro 0813 (max)1441
17GPT-5.6 Terra (max)1336GPT-5.6 Terra (max)1432
18K2-Horizon 375B A23B1303Gemini 3.8 Flash (high)1412
19DeepSeek V4 Pro 0813 (max)1261Qwen3.8 27B (xhigh)1409
20Gemini 3.8 Flash (high)1202K2-Horizon 375B A23B1349
21MiniMax-M31092MiniMax-M31230
22Nemotron 3 Ultra873Inkling1064
23Inkling829Nemotron 3 Ultra1000
24Gemini 3.5 Flash-Lite642Gemini 3.5 Flash-Lite970
25Mistral Medium 3.5515Muse Glimmer (high)774
26Muse Glimmer (high)469Mistral Medium 3.5747
27gpt-oss-120b (high)0gpt-oss-120b (high)596

The body also gives AA-Briefcase quality scores, and they cover only this pair:

MetricGrok 4.7Grok 4.6 (high)
Overall Elo16571546
Analytical quality Elo19941690
Presentation quality Elo14991519

The body does not separately repeat the reasoning tier for Grok 4.7 in this quality-score table. The overall Elo of 1657 in the same passage matches the 1657 for (xhigh) in the table above.

4. Output tokens, time, and speed

Output tokens per Intelligence Index question. The body uses “approximately”. The chart also prints a total and splits Answer and Reasoning into two segments. Adding 11k and 17k does not equal the 27k at the top of the bar. The table keeps the numbers printed on the chart and does not rewrite them into values that add up.

Model and tierApproximate figure in the bodyTotal on the chartAnswer on the chartReasoning on the chart
GPT-6 Astra (max)27k27k11k17k
Grok 4.6 (high)36k36k17k19k
Grok 4.6 (xhigh)38k38k18k20k
Muse Spark 1.3 (max)60k60k26k34k
Grok 4.7 (xhigh)81k81k22k59k

The body says 81k is 125% more than 36k and 196% more than 27k.

MetricValueDefinition
Answer output speedabout 188 tokens/secondBody: “for long prompts”. Prompt length is not stated
Time per Intelligence Index questionabout 7.1 minutesTime chart: weighted-average decode time, excluding TTFT and overhead

5. AA-Omniscience

MetricGrok 4.7 (xhigh)Grok 4.6 (high)
Index3230
Accuracy47%48%
Hallucination Rate29%34%

6. Model specifications

ItemWording in the article
Context window500k tokens, the same as Grok 4.6
Input / output price$2 / $6 per million tokens
Cache hit$0.50 per million tokens
Reasoning-effort rangelow to xhigh
Tier used in this evaluationxhigh
Dollar cost per evaluation questionnot stated

Conclusions and limitations

In this article, Grok 4.7 (xhigh)’s advantages concentrate on two kinds of tasks: long-horizon knowledge work under the unified harness (AA-Briefcase 1657, GDPval-AA 1695), and the coding-agent index on Grok Build (56). Analytical quality Elo is clearly higher than Grok 4.6 (high), and presentation quality Elo is lower. Under the unified harness, Terminal-Bench 4.0 and GDP.pdf rise slightly, and AA-LCR and AutomationBench-AA fall slightly. The body does not give absolute scores for these changes.

Token use is part of the limitation. About 81k output tokens is higher than the figures the body names: 36k for Grok 4.6 (high), 38k for Grok 4.6 (xhigh), 60k for Muse Spark 1.3 (max), and 27k for GPT-6 Astra (max). Of the 81k shown for Grok 4.7 on the chart, the Reasoning segment is 59k and the Answer segment is 22k.

The following cannot be folded into the measurements above:

  • Self-reported scores on the xAI release page, such as CursorBench, DeepSWE, and EEBench. This note did not open that page or cite its numbers. If DeepSWE or Terminal-Bench is cited, only this article’s Artificial Analysis / Grok Build definitions may be used.

  • The model index page. The article links to that page; this note does not use it to fill in scores, provider, or price.

  • high and xhigh. The +2 on the Intelligence Index chart corresponds to 44 for Grok 4.6 (xhigh). The AA-Briefcase, GDPval, and Omniscience comparisons are against Grok 4.6 (high). The coding-agent comparison is against Grok 4.6 (xhigh).

  • The two Terminal-Bench 4.0 results. The unified harness gives only +4.5 percentage points relative to Grok 4.6 (high). Grok Build gives 18% to 33%.

  • 7.1 minutes. This is decode time excluding TTFT and overhead.

  • $2 / $6 and cache $0.50. These are list prices. The article does not give the dollar cost of running the Intelligence Index or the coding index.

  • “top 4 AI labs”. This is headline wording. The article does not provide a lab ranking table.

  • Sample size, confidence-interval figures, temperature, seed, provider, API model ID, and the meaning of “with fallback” are all unspecified.

Reproduction notes

  1. Open https://artificialanalysis.ai/articles/benchmarking-grok-4-7 and check the date September 21, 2026 and the index of 46 in the headline.

  2. Read the body first, and keep xhigh, Grok Build, and the unified harness separate. Do not use the model index page to fill in fields this article does not have.

  3. Index scores, Elo, tokens, and coding-agent scores are in the images inside the article. Read the numbers at the tops of the bars and the angled model names, and keep the reasoning tiers printed on the chart. For the Elo charts, restate only the point estimates.

  4. On the large breakdown chart, the AA-Briefcase and GDPval axes are (Elo-500)/2000. Do not treat those percentages as raw Elo, and do not treat them as the 1657 or 1695 in the body.

  5. The article does not publish a question list, scoring script, random seed, or request parameters, so a third party cannot rerun it question by question from this article. What can be checked is the index versions, harness names, tiers, Elo scores, token counts, times, and the comparison-model set on this 2026-09-22 version of the page.

What this supports

  • Comparing Grok 4.7 (xhigh) on the composite Intelligence Index under Artificial Analysis’s unified harness, and on the Coding Agent Index in Grok Build, its first-party coding agent; judging long-horizon knowledge work such as AA-Briefcase and GDPval-AA, plus output tokens, decode time, and hallucination rate.

What this does not support

  • Substituting the scores on this page for the self-reported benchmarks on the xAI release page; treating Grok 4.6 high and xhigh as the same tier; treating Terminal-Bench 4.0 in the Coding Agent Index as the same-named item in the Intelligence Index; reading about 7.1 minutes as end-to-end wall-clock time that includes。

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis (the author type in the page JSON-LD is Organization) · Original publication date 2026-09-21 · Site edit date 2026-09-22

Open original source

Grok 4.7

Compare Grok 4.7 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Full review · English

Grok 4.7 Review: Same $2/$6 Price, About Twice the Tokens

A Grok 4.7 review of the unchanged $2/$6 rates, the jump to about 81k output tokens, and which workloads justify the extra work.

Pricing · English

Grok 4.7 Pricing: The $2/$6 Card and the Real Bill

Grok 4.7 keeps Grok 4.6's $2, $0.50, and $6 API rates. Effort, the 200k cliff, Fast, and Cursor's 256k line decide the bill.

Comparison · English

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

Related reviews

Arena: Grok 4.7 Has No Score Yet on the Agent Arena Net Improvement ChartAs of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.xAI Official Model Card: Grok 4.7 Safety Evaluation and Use BoundariesThe official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.xAI Official Release: Grok 4.7 Benchmark Scores and Capability PositioningOn the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.X: Artificial Analysis’s AA-Briefcase Chart — Grok 4.7 (xhigh) Composite Elo 1657On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).Grok 4.7 API setup on OpenRouterThe OpenRouter model page labels x-ai/grok-4.7 as SpaceXAI's Grok 4.7, lists input / output prices of $1.60 / $4.80 per million tokens, and gives OpenRouter SDK and cURL examples; the reasoning-level field in the request body, and the list prices for Low, Medium, and High, were not read in this collection.xAI Official Documentation: Grok 4.7 API Parameters and Reasoning LevelsThe model name on the public xAI API is grok-4.7. The reasoning levels listed in the documentation are Low, medium, high (default), or xhigh, with an input price of $2.00 / 1M tokens and an output price of $6.00 / 1M tokens. Grok 4.7 Fast is written as a faster deployment of the same model, billed at twice the standard token price, and it appears only in Cursor and Grok Build.Box's prompt and result for reviewing the Merewick claim with Grok 4.7 in AI StudioThis Box post is a 41-second Box Agent preview. In Box AI Studio, Grok 4.7 is selected and a claim-review prompt is entered for the folder "Active Commercial Property Claims", producing the file Claim Reconciliation Review Merewick Coastal Foods ASC-26-0184.md.