On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
xAI News (the browser title suffix is SpaceXAI; the footer copyright is SpaceXAI LLC) · Read evidenceGrok 4.7 · Reviews and evidence
Which Grok 4.7 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
xAI / SpaceXAI official model-card PDF (hosted on media.x.ai) · Read evidenceArtificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
Artificial Analysis · Read evidenceFull reviews and related reading
Selected evidence
xAI Official Release: Grok 4.7 Benchmark Scores and Capability Positioning
On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
- Source-specific observation
- On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
- Published conditions
- Treating the scores on this page as a third-party rerun; replacing the speed and price slogans in the headline with a measured comparison experiment; folding results for the fast variant, Grok 4.6, or other reasoning tiers into Grok 4.7 xHigh; comparing win rates or cost efficiency without a sample size, a harness。
xAI Official Model Card: Grok 4.7 Safety Evaluation and Use Boundaries
The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
- Source-specific observation
- The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
- Published conditions
- Treating official scores as a third-party rerun; applying a score from one tier to another; counting results for Grok 4.6, Grok 4.5, or comparison models as Grok 4.7 results; using unprotected capability probes as a substitute for post-deployment refusal behavior; or adding fast variants, consumer products。
Artificial Analysis: Grok 4.7 Intelligence Index and Coding Agent
Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
- Source-specific observation
- Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
- Published conditions
- Substituting the scores on this page for the self-reported benchmarks on the xAI release page; treating Grok 4.6 high and xhigh as the same tier; treating Terminal-Bench 4.0 in the Coding Agent Index as the same-named item in the Intelligence Index; reading about 7.1 minutes as end-to-end wall-clock time that includes。
X: Artificial Analysis’s AA-Briefcase Chart — Grok 4.7 (xhigh) Composite Elo 1657
On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).
- Source-specific observation
- On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).
- Published conditions
- Coding, multimodal work, general conversation, and any benchmark that does not appear on this chart or in this body text. The composite Elo on the chart cannot be substituted for a score from another harness.
All sources
All sources
xAI Official Release: Grok 4.7 Benchmark Scores and Capability Positioning
On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
- Source-specific observation
- On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
- Published conditions
- Treating the scores on this page as a third-party rerun; replacing the speed and price slogans in the headline with a measured comparison experiment; folding results for the fast variant, Grok 4.6, or other reasoning tiers into Grok 4.7 xHigh; comparing win rates or cost efficiency without a sample size, a harness。
xAI Official Model Card: Grok 4.7 Safety Evaluation and Use Boundaries
The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
- Source-specific observation
- The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
- Published conditions
- Treating official scores as a third-party rerun; applying a score from one tier to another; counting results for Grok 4.6, Grok 4.5, or comparison models as Grok 4.7 results; using unprotected capability probes as a substitute for post-deployment refusal behavior; or adding fast variants, consumer products。
Artificial Analysis: Grok 4.7 Intelligence Index and Coding Agent
Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
- Source-specific observation
- Artificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.
- Published conditions
- Substituting the scores on this page for the self-reported benchmarks on the xAI release page; treating Grok 4.6 high and xhigh as the same tier; treating Terminal-Bench 4.0 in the Coding Agent Index as the same-named item in the Intelligence Index; reading about 7.1 minutes as end-to-end wall-clock time that includes。
X: Artificial Analysis’s AA-Briefcase Chart — Grok 4.7 (xhigh) Composite Elo 1657
On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).
- Source-specific observation
- On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).
- Published conditions
- Coding, multimodal work, general conversation, and any benchmark that does not appear on this chart or in this body text. The composite Elo on the chart cannot be substituted for a score from another harness.
Arena: Grok 4.7 Has No Score Yet on the Agent Arena Net Improvement Chart
As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.
- Source-specific observation
- As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.
- Published conditions
- It cannot be used to judge Grok 4.7's actual net improvement, ranking, or blind-evaluation win rate, and it cannot be extended to quality on text, vision, code, or document tasks.
Same-Task Cursor Ultra Usage for Grok 4.7 and Grok 4.6
On Cursor Ultra, at extra high, with fast mode off, the author estimates from how fast the same task reduced plan usage that Grok 4.7 consumes about 2.5 times as much as Grok 4.6. This is a short timing on a single account and cannot be treated as a quality benchmark.
- Source-specific observation
- On Cursor Ultra, at extra high, with fast mode off, the author estimates from how fast the same task reduced plan usage that Grok 4.7 consumes about 2.5 times as much as Grok 4.6. This is a short timing on a single account and cannot be treated as a quality benchmark.
- Published conditions
- Whether the code is correct, answer quality, an official API bill in dollars, fast mode, other reasoning tiers, other clients, and any work other than the task the author did not write down.
Watching Grok 4.6 and Grok 4.7 Age of Empires Footage Side by Side on Reddit
This is a roughly 55-second, silent, left-right split-screen video: the left side is labeled Grok 4.6, the right side is labeled Grok 4.7, and both show Age of Empires II-style towns and combat. The post gives no prompt, client, reasoning tier, or win/loss rules, so it can only be treated as a side-by-side impression.
- Source-specific observation
- This is a roughly 55-second, silent, left-right split-screen video: the left side is labeled Grok 4.6, the right side is labeled Grok 4.7, and both show Age of Empires II-style towns and combat. The post gives no prompt, client, reasoning tier, or win/loss rules, so it can only be treated as a side-by-side impression.
- Published conditions
- Who finished building, whether the two sides are the same map or the same prompt, play quality, official API cost, and other games or clients.
A Usage Report on Grok 4.7 Handling an Insurance Claim in Box AI Studio
In a video post, the author says that while Grok 4.7 handled a $2 million insurance claim in Box AI Studio, it pointed out an $82,000 duplicate invoice and a missing $64,000 supplier credit, and cited the claims review for the adjuster. Collection read only that sentence; the video frames were not transcribed.
- Source-specific observation
- In a video post, the author says that while Grok 4.7 handled a $2 million insurance claim in Box AI Studio, it pointed out an $82,000 duplicate invoice and a missing $64,000 supplier credit, and cited the claims review for the adjuster. Collection read only that sentence; the video frames were not transcribed.
- Published conditions
- Whether the claim was paid, the net amount saved, other lines of insurance or other amounts, other clients, official benchmarks, API cost, and operating steps in the video that were never written out as text.
r/cursor users' first impressions of Grok 4.7
In the 7 comments within hours of the post, four gave subjective impressions of Grok 4.7: not as slow as 4.6, feeling like an opus5 that does not overthink as much, writing that seems better but is still being tested, and Devin's SWE 2.0 plus fusion api described as better. Nobody wrote down a task, a timing, or a score.
- Source-specific observation
- In the 7 comments within hours of the post, four gave subjective impressions of Grok 4.7: not as slow as 4.6, feeling like an opus5 that does not overthink as much, writing that seems better but is still being tested, and Devin's SWE 2.0 plus fusion api described as better. Nobody wrote down a task, a timing, or a score.
- Published conditions
- Whether the code is correct, benchmark rank, price, plan usage, an official API bill, and any comparison that needs a fixed task and scoring rules.
X: eric zakariasson Side-by-Side Demo of Grok 4.6 and Grok 4.7 Building Age of Empires II
This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.
- Source-specific observation
- This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.
- Published conditions
- Which side won, whether the two sides used the same prompt or the same map, play quality, coding and other games, and benchmarks the quoted image does not list.
Grok 4.7
Compare Grok 4.7 in Tabbit
Model access, features, and permissions depend on your current client account.