Grok 4.7

Grok 4.7 · Reviews and evidence

Which Grok 4.7 conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.

Artificial Analysis · Read evidence

Full reviews and related reading

Read the full analysis

Full review · English

Grok 4.7 Review: Same $2/$6 Price, About Twice the Tokens

A Grok 4.7 review of the unchanged $2/$6 rates, the jump to about 81k output tokens, and which workloads justify the extra work.

Pricing · English

Grok 4.7 Pricing: The $2/$6 Card and the Real Bill

Grok 4.7 keeps Grok 4.6's $2, $0.50, and $6 API rates. Effort, the 200k cliff, Fast, and Cursor's 256k line decide the bill.

Selected evidence

OfficialVendor report

xAI Official Release: Grok 4.7 Benchmark Scores and Capability Positioning

On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.

SourcexAI News (the browser title suffix is SpaceXAI; the footer copyright is SpaceXAI LLC)
Published2026-09-21
Collected2026-09-22
Source-specific observation
On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
Published conditions
Treating the scores on this page as a third-party rerun; replacing the speed and price slogans in the headline with a measured comparison experiment; folding results for the fast variant, Grok 4.6, or other reasoning tiers into Grok 4.7 xHigh; comparing win rates or cost efficiency without a sample size, a harness。
Capability
Media / benchmarkVendor report

xAI Official Model Card: Grok 4.7 Safety Evaluation and Use Boundaries

The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.

SourcexAI / SpaceXAI official model-card PDF (hosted on media.x.ai)
Published2026-09-21
Collected2026-09-22
Source-specific observation
The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
Published conditions
Treating official scores as a third-party rerun; applying a score from one tier to another; counting results for Grok 4.6, Grok 4.5, or comparison models as Grok 4.7 results; using unprotected capability probes as a substitute for post-deployment refusal behavior; or adding fast variants, consumer products。
Capability
Media / benchmarkIndependent measurement

Artificial Analysis: Grok 4.7 Intelligence Index and Coding Agent

Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.

SourceArtificial Analysis
Published2026-09-21
Collected2026-09-22
Source-specific observation
Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
Published conditions
Substituting the scores on this page for the self-reported benchmarks on the xAI release page; treating Grok 4.6 high and xhigh as the same tier; treating Terminal-Bench 4.0 in the Coding Agent Index as the same-named item in the Intelligence Index; reading about 7.1 minutes as end-to-end wall-clock time that includes。
AgentCoding
CommunityIndependent measurement

X: Artificial Analysis’s AA-Briefcase Chart — Grok 4.7 (xhigh) Composite Elo 1657

On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).

SourceX
Published2026-09-22
Collected2026-09-22
Source-specific observation
On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).
Published conditions
Coding, multimodal work, general conversation, and any benchmark that does not appear on this chart or in this body text. The composite Elo on the chart cannot be substituted for a score from another harness.
Capability

All sources

All sources

10 / 10
OfficialVendor report

xAI Official Release: Grok 4.7 Benchmark Scores and Capability Positioning

On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.

SourcexAI News (the browser title suffix is SpaceXAI; the footer copyright is SpaceXAI LLC)
Published2026-09-21
Collected2026-09-22
Source-specific observation
On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
Published conditions
Treating the scores on this page as a third-party rerun; replacing the speed and price slogans in the headline with a measured comparison experiment; folding results for the fast variant, Grok 4.6, or other reasoning tiers into Grok 4.7 xHigh; comparing win rates or cost efficiency without a sample size, a harness。
Capability
Media / benchmarkVendor report

xAI Official Model Card: Grok 4.7 Safety Evaluation and Use Boundaries

The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.

SourcexAI / SpaceXAI official model-card PDF (hosted on media.x.ai)
Published2026-09-21
Collected2026-09-22
Source-specific observation
The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
Published conditions
Treating official scores as a third-party rerun; applying a score from one tier to another; counting results for Grok 4.6, Grok 4.5, or comparison models as Grok 4.7 results; using unprotected capability probes as a substitute for post-deployment refusal behavior; or adding fast variants, consumer products。
Capability
Media / benchmarkIndependent measurement

Artificial Analysis: Grok 4.7 Intelligence Index and Coding Agent

Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.

SourceArtificial Analysis
Published2026-09-21
Collected2026-09-22
Source-specific observation
Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
Published conditions
Substituting the scores on this page for the self-reported benchmarks on the xAI release page; treating Grok 4.6 high and xhigh as the same tier; treating Terminal-Bench 4.0 in the Coding Agent Index as the same-named item in the Intelligence Index; reading about 7.1 minutes as end-to-end wall-clock time that includes。
AgentCoding
CommunityIndependent measurement

X: Artificial Analysis’s AA-Briefcase Chart — Grok 4.7 (xhigh) Composite Elo 1657

On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).

SourceX
Published2026-09-22
Collected2026-09-22
Source-specific observation
On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).
Published conditions
Coding, multimodal work, general conversation, and any benchmark that does not appear on this chart or in this body text. The composite Elo on the chart cannot be substituted for a score from another harness.
Capability
CommunityPlatform telemetry

Arena: Grok 4.7 Has No Score Yet on the Agent Arena Net Improvement Chart

As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.

SourceX (@arena)
Published2026-09-21
Collected2026-09-22
Source-specific observation
As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.
Published conditions
It cannot be used to judge Grok 4.7's actual net improvement, ranking, or blind-evaluation win rate, and it cannot be extended to quality on text, vision, code, or document tasks.
AgentCapability
CommunityEditorial analysis

Same-Task Cursor Ultra Usage for Grok 4.7 and Grok 4.6

On Cursor Ultra, at extra high, with fast mode off, the author estimates from how fast the same task reduced plan usage that Grok 4.7 consumes about 2.5 times as much as Grok 4.6. This is a short timing on a single account and cannot be treated as a quality benchmark.

SourceReddit (r/cursor)
Published2026-09-21
Collected2026-09-22
Source-specific observation
On Cursor Ultra, at extra high, with fast mode off, the author estimates from how fast the same task reduced plan usage that Grok 4.7 consumes about 2.5 times as much as Grok 4.6. This is a short timing on a single account and cannot be treated as a quality benchmark.
Published conditions
Whether the code is correct, answer quality, an official API bill in dollars, fast mode, other reasoning tiers, other clients, and any work other than the task the author did not write down.
Capability
CommunityEditorial analysis

Watching Grok 4.6 and Grok 4.7 Age of Empires Footage Side by Side on Reddit

This is a roughly 55-second, silent, left-right split-screen video: the left side is labeled Grok 4.6, the right side is labeled Grok 4.7, and both show Age of Empires II-style towns and combat. The post gives no prompt, client, reasoning tier, or win/loss rules, so it can only be treated as a side-by-side impression.

SourceReddit (r/xai)
Published2026-09-21
Collected2026-09-22
Source-specific observation
This is a roughly 55-second, silent, left-right split-screen video: the left side is labeled Grok 4.6, the right side is labeled Grok 4.7, and both show Age of Empires II-style towns and combat. The post gives no prompt, client, reasoning tier, or win/loss rules, so it can only be treated as a side-by-side impression.
Published conditions
Who finished building, whether the two sides are the same map or the same prompt, play quality, official API cost, and other games or clients.
Capability
CommunityPersonal experience

A Usage Report on Grok 4.7 Handling an Insurance Claim in Box AI Studio

In a video post, the author says that while Grok 4.7 handled a $2 million insurance claim in Box AI Studio, it pointed out an $82,000 duplicate invoice and a missing $64,000 supplier credit, and cited the claims review for the adjuster. Collection read only that sentence; the video frames were not transcribed.

SourceReddit (r/xai)
Published2026-09-21
Collected2026-09-22
Source-specific observation
In a video post, the author says that while Grok 4.7 handled a $2 million insurance claim in Box AI Studio, it pointed out an $82,000 duplicate invoice and a missing $64,000 supplier credit, and cited the claims review for the adjuster. Collection read only that sentence; the video frames were not transcribed.
Published conditions
Whether the claim was paid, the net amount saved, other lines of insurance or other amounts, other clients, official benchmarks, API cost, and operating steps in the video that were never written out as text.
ResearchCapability
CommunityPersonal experience

r/cursor users' first impressions of Grok 4.7

In the 7 comments within hours of the post, four gave subjective impressions of Grok 4.7: not as slow as 4.6, feeling like an opus5 that does not overthink as much, writing that seems better but is still being tested, and Devin's SWE 2.0 plus fusion api described as better. Nobody wrote down a task, a timing, or a score.

SourceReddit (r/cursor)
Published2026-09-21
Collected2026-09-22
Source-specific observation
In the 7 comments within hours of the post, four gave subjective impressions of Grok 4.7: not as slow as 4.6, feeling like an opus5 that does not overthink as much, writing that seems better but is still being tested, and Devin's SWE 2.0 plus fusion api described as better. Nobody wrote down a task, a timing, or a score.
Published conditions
Whether the code is correct, benchmark rank, price, plan usage, an official API bill, and any comparison that needs a fixed task and scoring rules.
Capability
CommunityEditorial analysis

X: eric zakariasson Side-by-Side Demo of Grok 4.6 and Grok 4.7 Building Age of Empires II

This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.

SourceX
Published2026-09-21
Collected2026-09-22
Source-specific observation
This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.
Published conditions
Which side won, whether the two sides used the same prompt or the same map, play quality, coding and other games, and benchmarks the quoted image does not list.
Capability

Grok 4.7

Compare Grok 4.7 in Tabbit

Model access, features, and permissions depend on your current client account.