Grok 4.7 · Community source · Editorial analysis
On Cursor Ultra, at extra high, with fast mode off, the author estimates from how fast the same task reduced plan usage that Grok 4.7 consumes about 2.5 times as much as Grok 4.6. This is a short timing on a single account and cannot be treated as a quality benchmark.
On Cursor Ultra, at extra high, with fast mode off, the author estimates from how fast the same task reduced plan usage that Grok 4.7 consumes about 2.5 times as much as Grok 4.6. This is a short timing on a single account and cannot be treated as a quality benchmark.
Tasks suitable for a judgment: On the same Cursor Ultra account, with extra high and fast mode off, how fast Grok 4.7 burns plan percentage relative to Grok 4.6.
Tasks unsuitable for extrapolation: Whether the code is correct, answer quality, an official API bill in dollars, fast mode, other reasoning tiers, other clients, and any work other than the task the author did not write down.
Applicable model versions: Grok 4.7 and Grok 4.6 in the post body. The points in the attached figure are labeled Grok Build’s Grok 4.6 (xhigh) and Grok 4.7 (xhigh), a different measurement from the Cursor usage.
Test environment or client: Cursor Ultra. The repository, the task text, and the client version are not stated.
Reasoning tiers and parameters: extra high effort, with fast mode off for both versions. Temperature and other sampling parameters are not stated.
The author writes that they ran the same task for two consecutive days, and about 3 hours before posting, at 22:00 Central European Time, switched from Grok 4.6 to Grok 4.7. The comparison is the time it takes Cursor usage to fall by 1 percentage point. The post also includes an Artificial Analysis scatter plot. The author uses the average dollar cost per task on that chart as a side-by-side check on their own 2.5-times judgment. The post does not give the task input, the output, or the scoring rules.
The intervals the author wrote down are:
Grok 4.7: 82%→81% was 34 minutes; 81%→80% was 55 minutes, or 71 minutes if the 15-minute break in the middle is counted; 80%→79% was 38 minutes.
Grok 4.6: the author says this was the same task earlier the same day; 85%→84% was 133 minutes, and 84%→83% was 101 minutes.
From the author’s two sets of numbers, after excluding the break, 4.6’s average interval is about 117 minutes and 4.7’s is about 42 minutes, a ratio of about 2.8. If the middle 4.7 interval is counted as 71 minutes of wall-clock time, the ratio is about 2.5. The author summarizes the result as about 2.5 times, and judges the gap too large relative to a benchmark improvement of about 10%, so most tasks would switch back to 4.6.
The clocks on the usage chart match 4.7’s 34 minutes, 71 minutes, and 38 minutes, and 4.6’s 101 minutes. The visible interval from 85% (17:52) to 84% (19:35) is about 103 minutes, which does not match the 133 minutes in the body. The span closer to 133 minutes is from 86% (15:41) to 85% (17:52), about 131 minutes.
The author also writes that the average cost per task in the chart is $8.82 for Grok 4.7 and $3.53 for Grok 4.6, and that this ratio is also 2.5 times. Those dollar figures are the author’s retelling of the attached figure.
Short quotations from the body: “Both 4.7 and 4.6 were used on extra high effort and both not on fast mode.” and “Grok 4.7 is about 2.5 times as expensive as 4.6”.
The first image is a dark usage log. The overlay lists from 88% (01:09) down to 79% (00:50), and the last row reads gem. 63 min / 1%. Differences between adjacent clocks are about: 88%→87% 6 minutes, 87%→86% 14 hours 26 minutes, 86%→85% 131 minutes, 85%→84% 103 minutes, 84%→83% 101 minutes, 83%→82% 71 minutes, 82%→81% 34 minutes, 81%→80% 71 minutes, 80%→79% 38 minutes. The long gap from 87% to 86% shows that this table contains idle time, and 63 min / 1% is a mixed average. The overlay covers part of the table behind it. The background shows rows for Grok, Codex, Claude, and Cursor, plus the Dutch dates Woensdag 23-09, Donderdag 24-09, and Zaterdag 26-09. The Cursor row at the bottom also has Zaterdag 17-10 00:18. These dates do not line up with the posting time of 2026-09-21, and the figure does not mark which stretch belongs to 4.7. At collection time, the Cursor bar on the page showed 79% remaining.
The second figure is titled “Artificial Analysis Coding Agent Index vs. Cost per Task”. The subtitle states that the horizontal axis is the average token-billed API cost per task in dollars, and the vertical axis is the Coding Agent Index. The two legible Grok labels on the chart are Grok Build - Grok 4.6 (xhigh) (SpaceXAI) and Grok Build - Grok 4.7 (xhigh) (SpaceXAI). The 4.7 point is higher than the 4.6 point and farther to the right. $8.82 and $3.53 are not printed directly on the axis ticks. The same chart includes points from other labs; they are not rewritten here as a ranking.
At collection time the post had about 29 points, 4 comments, and was labeled Question / Discussion. The gist of the comments:
KeenAsGreen calls it still a “meme llm”, with no task or data.
truecakesnake thinks this release is poor, that benchmarks no longer explain the issue, and that the price is clearly too high.
Vegetable-Piano-5313 thinks the cost of standard-tier Grok 4.7 is roughly equivalent to Grok 4.6 fast mode.
Dynamix86 replies with the short quotation: “25% more expensive than 4.6 fast mode even.” That reply does not include a fast-mode timing.
The usable conclusion is limited to this author’s Cursor Ultra account: at extra high, with fast mode off, the few percentage points after the switch to Grok 4.7 dropped faster than the earlier same-day Grok 4.6 stretch the author identified. The author therefore plans to go back to 4.6 for most tasks.
Cursor percentage time and Artificial Analysis’s dollar cost per task use different clients, billing units, and task sets. The attached figure is the xhigh API cost for Grok Build; the body measures Cursor plan percentage. The roughly 10% benchmark improvement the author mentions is an impression of that chart. The post has no sample size, task list, or scoring script.
The task content is not stated. The 133 minutes does not match the clock difference for 85%→84% in the figure. Being 25% more expensive than fast mode appears only in one reply. The first two comments are attitudes and cannot be recorded as benchmark results.
Opening the permalink above shows the body, both figures, and the 4 comments present at collection time. Redoing the author’s own usage comparison requires Cursor Ultra, the same unpublished task, Grok 4.6 and Grok 4.7, extra high, and fast mode off, then a record of the percentage and the time. The post has no prompt and no output, so the quality part cannot be reproduced. $8.82 and $3.53 are the author’s readings of the screenshot. At collection time, the original Artificial Analysis page was not opened separately to check them.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Reddit (r/cursor) · Dynamix86 · Original publication date 2026-09-21 · Site edit date 2026-09-22
Open original sourceGrok 4.7
Download the Tabbit client to check model access