Grok 4.7 · Community source · Editorial analysis
This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.
This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.
Tasks suitable for a judgment: In this video, which age the footage under each label is sitting in, and roughly what the top-bar numbers are at the sampled timestamps.
Tasks unsuitable for extrapolation: Which side won, whether the two sides used the same prompt or the same map, play quality, coding and other games, and benchmarks the quoted image does not list.
Applicable model versions: The body and the on-screen labels say Grok 4.6 and Grok 4.7. Finer checkpoints and API model IDs are not stated.
Test environment or client: The body mentions Cursor, Grok Build, and the API, but does not say which of those this video used. Not stated.
Reasoning tiers and parameters: The video does not state them. Column headers on the quoted image label Grok 4.7 as xHigh and Grok 4.6 as High.
The author did not write steps, inputs, or scoring rules. The comparison comes from the split-screen labels on the embedded video. The video is 55.96 seconds long at 1920×720, and the player progress runs to 0:56. The player indicates that this video has no captions. During collection, playback was paused at 0:00, 0:08, 0:16, 0:24, 0:32, 0:40, 0:48, and 0:54.5, and the top bar and bottom labels were read from the native frame. The leftmost digit on the left side’s top bar is cropped and unstable from frame to frame, so it is not included in the table below.
The same permalink also quotes a post by @SpaceXAI with a four-column table. The numbers in that table come from the image, not from scores on the gameplay footage.
Body text after clicking “Show original”:
grok 4.7 is here, and its our best model so far! try it out in cursor, grok build, api or anywhere you get your… This is a necessary excerpt; read the original source for full context.
“The best model so far” is the author’s own wording. The video does not give a ranking, a countdown, or a win/loss caption.
Top bar as read from the sampled frames:
| Time | Left: Grok 4.6 | Right: Grok 4.7 |
|---|---|---|
| 0:00, 0:08, 0:16, 0:24 | 95, 85; population 46/155; Age II Idle 34 | 45, 35, 25, 40; Age 1 |
| 0:32 | 95, 85; population 46/155; Age II Idle 34 | 65, 35, 25, 50; Age 1 |
| 0:40 | 90, 40; population 45/155; Age II Idle 33 | 65, 35, 25, 50; Age 1 |
| 0:48 | 90, 40; population 46/155; Age II Idle 34 | 115, 35, 55, 50; Age 1 |
| 0:54.5 | 90, 40; population 46/155; Age II Idle 22 | 115, 35, 55, 50; Age 1 |
At 0:00 and 0:08, the left side is a coastal town, with large yellow text TIDEHOLD in the center and The Brinewatch Landing on the line below. From 0:16 on, the left camera leaves that place-name banner; a waterwheel, docks, towers, and soldiers are visible. On the right, 0:00 and 0:08 show a stone castle town; 0:16 and 0:24 cut to a construction site with scaffolding; 0:32 shows buildings by the water; 0:48 and 0:54.5 show a large melee. In these sampled frames, the age button on the right always reads Age 1, and there is no place-name banner like the one on the left.
On these frames, the command bar at the bottom left reads Dock, House, and Barracks. The same three words were not read on the bottom right. The two sides do not share a scoreboard.
At collection time, this post’s engagement counts were 153 replies, 99 reposts, 1682 likes, 168 bookmarks, and 2723271 views. These are page counts, not model scores.
Visible English on the quoted post:
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
That quote’s datetime is 2026-09-21T16:17:53.000Z, and its permalink is https://x.com/SpaceXAI/status/2102069815225586149. The four columns in the image, left to right, are Grok 4.7 xHigh, Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max. The footnote reads *High Effort and is attached only to Grok 4.7’s DeepSWE cell. Sample size, provider, run date, and prompt are not stated.
| Row | Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| Input token price, $ per million | $2 | $2 | $4 | $10 |
| Output token price, $ per million | $6 | $6 | $20 | $50 |
| Software engineering, CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| Software engineering, DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| Electrical engineering, EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| Multi-Hour office work, AA-Briefcase v1.1 | 1657 | 1546 | 1487 | 1678 |
| Multi-Hour terminal work, Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Legal work, Harvey Legal Agent | 19.6% | 15.8% | 2.5% | 6.7% |
| Clinical reasoning, HealthBench Pro | 56.7% | 48.5% | 60.5% | 62.1% |
This video only shows, in one side-by-side view, what the labels, the age text, and the sampled top bars are on each side. The right side sitting on Age 1 and the left side sitting on Age II are words on the screen. The post does not declare a winner from them, and they cannot be recast as “4.7 failed to advance” or “4.6 plays the game better.”
In the quoted image, 1657 is Grok 4.7 xHigh on the AA-Briefcase v1.1 row, and the comparison column is Grok 4.6 High at 1546. That figure is not the game video’s score, and it cannot be subtracted directly from a Grok 4.6 number at a different tier. The table does not name an overall winner.
On 2026-09-22 the permalink above was opened. The page was logged in, and no captcha appeared. After clicking “Show original,” the English body was transcribed, the embedded video was played, and it was paused at the eight timestamps above to read the top bar of the 1920×720 frame. The quoted image was downloaded from the in-post image URL and the table was read from it. The English of the quoted lines was transcribed in the same task by opening that quote’s permalink and clicking “Show original” again. The post does not give a prompt, client, reasoning tier, or original run log, so this demo cannot be rerun from the post alone.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X · eric zakariasson (@ericzakariasson, verified account) · Original publication date 2026-09-21 · Site edit date 2026-09-22
Open original sourceGrok 4.7
Download the Tabbit client to check model access