Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Grok 4.7 · Community source · Editorial analysis

X: eric zakariasson Side-by-Side Demo of Grok 4.6 and Grok 4.7 Building Age of Empires II

This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.

Community sourceEditorial analysisEdited 2026-09-22

Test conditions

Source-specific observation
This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.
Published conditions
Which side won, whether the two sides used the same prompt or the same map, play quality, coding and other games, and benchmarks the quoted image does not list.

Key data and applicable tasks

One-sentence takeaway

This post uses a roughly 56-second, 1920×720 side-by-side video to show Grok 4.6 and Grok 4.7 building Age of Empires II. The sampled frames show the labels, ages, and top-bar numbers on both sides, but no prompt, reasoning tier, or win/loss caption. The quoted @SpaceXAI image is a separate price and benchmark table; that table is not the score for this gameplay footage.

Use cases

  • Tasks suitable for a judgment: In this video, which age the footage under each label is sitting in, and roughly what the top-bar numbers are at the sampled timestamps.

  • Tasks unsuitable for extrapolation: Which side won, whether the two sides used the same prompt or the same map, play quality, coding and other games, and benchmarks the quoted image does not list.

  • Applicable model versions: The body and the on-screen labels say Grok 4.6 and Grok 4.7. Finer checkpoints and API model IDs are not stated.

  • Test environment or client: The body mentions Cursor, Grok Build, and the API, but does not say which of those this video used. Not stated.

  • Reasoning tiers and parameters: The video does not state them. Column headers on the quoted image label Grok 4.7 as xHigh and Grok 4.6 as High.

Evaluation method

The author did not write steps, inputs, or scoring rules. The comparison comes from the split-screen labels on the embedded video. The video is 55.96 seconds long at 1920×720, and the player progress runs to 0:56. The player indicates that this video has no captions. During collection, playback was paused at 0:00, 0:08, 0:16, 0:24, 0:32, 0:40, 0:48, and 0:54.5, and the top bar and bottom labels were read from the native frame. The leftmost digit on the left side’s top bar is cropped and unstable from frame to frame, so it is not included in the table below.

The same permalink also quotes a post by @SpaceXAI with a four-column table. The numbers in that table come from the image, not from scores on the gameplay footage.

Key results

Body text after clicking “Show original”:

grok 4.7 is here, and its our best model so far! try it out in cursor, grok build, api or anywhere you get your… This is a necessary excerpt; read the original source for full context.

“The best model so far” is the author’s own wording. The video does not give a ranking, a countdown, or a win/loss caption.

Top bar as read from the sampled frames:

TimeLeft: Grok 4.6Right: Grok 4.7
0:00, 0:08, 0:16, 0:2495, 85; population 46/155; Age II Idle 3445, 35, 25, 40; Age 1
0:3295, 85; population 46/155; Age II Idle 3465, 35, 25, 50; Age 1
0:4090, 40; population 45/155; Age II Idle 3365, 35, 25, 50; Age 1
0:4890, 40; population 46/155; Age II Idle 34115, 35, 55, 50; Age 1
0:54.590, 40; population 46/155; Age II Idle 22115, 35, 55, 50; Age 1

At 0:00 and 0:08, the left side is a coastal town, with large yellow text TIDEHOLD in the center and The Brinewatch Landing on the line below. From 0:16 on, the left camera leaves that place-name banner; a waterwheel, docks, towers, and soldiers are visible. On the right, 0:00 and 0:08 show a stone castle town; 0:16 and 0:24 cut to a construction site with scaffolding; 0:32 shows buildings by the water; 0:48 and 0:54.5 show a large melee. In these sampled frames, the age button on the right always reads Age 1, and there is no place-name banner like the one on the left.

On these frames, the command bar at the bottom left reads Dock, House, and Barracks. The same three words were not read on the bottom right. The two sides do not share a scoreboard.

Raw data

At collection time, this post’s engagement counts were 153 replies, 99 reposts, 1682 likes, 168 bookmarks, and 2723271 views. These are page counts, not model scores.

Visible English on the quoted post:

Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.

That quote’s datetime is 2026-09-21T16:17:53.000Z, and its permalink is https://x.com/SpaceXAI/status/2102069815225586149. The four columns in the image, left to right, are Grok 4.7 xHigh, Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max. The footnote reads *High Effort and is attached only to Grok 4.7’s DeepSWE cell. Sample size, provider, run date, and prompt are not stated.

RowGrok 4.7 xHighGrok 4.6 HighGPT-5.6 Sol MaxFable 5.1 Max
Input token price, $ per million$2$2$4$10
Output token price, $ per million$6$6$20$50
Software engineering, CursorBench 4.046.3%40.4%41.7%51.8%
Software engineering, DeepSWE v1.171.0%*65.2%72.7%70.0%
Electrical engineering, EEBench64.0%53.0%39.4%56.4%
Multi-Hour office work, AA-Briefcase v1.11657154614871678
Multi-Hour terminal work, Terminal-Bench 4.038.0%20.3%37.3%57.9%
Legal work, Harvey Legal Agent19.6%15.8%2.5%6.7%
Clinical reasoning, HealthBench Pro56.7%48.5%60.5%62.1%

Conclusions and limitations

This video only shows, in one side-by-side view, what the labels, the age text, and the sampled top bars are on each side. The right side sitting on Age 1 and the left side sitting on Age II are words on the screen. The post does not declare a winner from them, and they cannot be recast as “4.7 failed to advance” or “4.6 plays the game better.”

In the quoted image, 1657 is Grok 4.7 xHigh on the AA-Briefcase v1.1 row, and the comparison column is Grok 4.6 High at 1546. That figure is not the game video’s score, and it cannot be subtracted directly from a Grok 4.6 number at a different tier. The table does not name an overall winner.

Reproduction notes

On 2026-09-22 the permalink above was opened. The page was logged in, and no captcha appeared. After clicking “Show original,” the English body was transcribed, the embedded video was played, and it was paused at the eight timestamps above to read the top bar of the 1920×720 frame. The quoted image was downloaded from the in-post image URL and the table was read from it. The English of the quoted lines was transcribed in the same task by opening that quote’s permalink and clicking “Show original” again. The post does not give a prompt, client, reasoning tier, or original run log, so this demo cannot be rerun from the post alone.

What this supports

  • In this video, which age the footage under each label is sitting in, and roughly what the top-bar numbers are at the sampled timestamps.

What this does not support

  • Which side won, whether the two sides used the same prompt or the same map, play quality, coding and other games, and benchmarks the quoted image does not list.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X · eric zakariasson (@ericzakariasson, verified account) · Original publication date 2026-09-21 · Site edit date 2026-09-22

Open original source

Grok 4.7

Compare Grok 4.7 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Full review · English

Grok 4.7 Review: Same $2/$6 Price, About Twice the Tokens

A Grok 4.7 review of the unchanged $2/$6 rates, the jump to about 81k output tokens, and which workloads justify the extra work.

Pricing · English

Grok 4.7 Pricing: The $2/$6 Card and the Real Bill

Grok 4.7 keeps Grok 4.6's $2, $0.50, and $6 API rates. Effort, the 200k cliff, Fast, and Cursor's 256k line decide the bill.

Comparison · English

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

Related reviews

X: Artificial Analysis’s AA-Briefcase Chart — Grok 4.7 (xhigh) Composite Elo 1657On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).xAI Official Release: Grok 4.7 Benchmark Scores and Capability PositioningOn the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.xAI Official Model Card: Grok 4.7 Safety Evaluation and Use BoundariesThe official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.Arena: Grok 4.7 Has No Score Yet on the Agent Arena Net Improvement ChartAs of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.Box's prompt and result for reviewing the Merewick claim with Grok 4.7 in AI StudioThis Box post is a 41-second Box Agent preview. In Box AI Studio, Grok 4.7 is selected and a claim-review prompt is entered for the folder "Active Commercial Property Claims", producing the file Claim Reconciliation Review Merewick Coastal Foods ASC-26-0184.md.Grok 4.7 API setup on OpenRouterThe OpenRouter model page labels x-ai/grok-4.7 as SpaceXAI's Grok 4.7, lists input / output prices of $1.60 / $4.80 per million tokens, and gives OpenRouter SDK and cURL examples; the reasoning-level field in the request body, and the list prices for Low, Medium, and High, were not read in this collection.xAI Official Documentation: Grok 4.7 API Parameters and Reasoning LevelsThe model name on the public xAI API is grok-4.7. The reasoning levels listed in the documentation are Low, medium, high (default), or xhigh, with an input price of $2.00 / 1M tokens and an output price of $6.00 / 1M tokens. Grok 4.7 Fast is written as a faster deployment of the same model, billed at twice the standard token price, and it appears only in Cursor and Grok Build.