Grok 4.7 · Community source · Platform telemetry
As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.
As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.
Tasks suitable for a judgment: Whether Arena has already published a net improvement score for Grok 4.7, and whether, at collection time, the community predicted 4.7 relative to the historical trend as Better, On trend, or Worse.
Tasks unsuitable for extrapolation: It cannot be used to judge Grok 4.7's actual net improvement, ranking, or blind-evaluation win rate, and it cannot be extended to quality on text, vision, code, or document tasks.
Applicable model version: The model discussed in the post that still has no score is Grok 4.7. The points already plotted belong to Grok 4.3 (High), Grok Build 0.1, Grok 4.5, and Grok 4.6 (xHigh), and must not be written as Grok 4.7 results.
Test environment or client: The chart title is Agent Arena, the source credit reads SOURCE: AGENT ARENA, and the footnote is RESULTS 7 DAYS AFTER RELEASE. Grok 4.7 has no corresponding result point on this chart.
Reasoning tiers and parameters: Not stated for Grok 4.7. The chart labels only the older versions Grok 4.3 (High) and Grok 4.6 (xHigh). Temperature, tool configuration, and sample size are not stated.
This is not a Grok 4.7 leaderboard row with a public vote count, and it is not a controlled evaluation of Grok 4.7. In the post, @arena shows a historical net-improvement line chart and starts a prediction poll in a reply on the same page. The original post text is: Grok 4.7 by @SpaceXAI just dropped. Looking at how past Grok versions have trended on Agent Arena's net improvement score, where do you think 4.7 lands? Poll below. Scores coming soon as you vote on @arena!
The vertical axis is NET IMPROVEMENT (%), with visible ticks at 10%, 5%, 0%, -5%, and -10%. The horizontal axis is labeled 8 Jun 2026, 13 Jul 2026, 27 Aug 2026, and Coming soon. The plotted points have vertical error bars, but the chart does not print values for those bars. The footnote reads RESULTS 7 DAYS AFTER RELEASE. It does not state the task set, the prompts, the sample size, or the vote counts that produced these historical scores.
The prediction poll is the next @arena reply on the same page: https://x.com/arena/status/2102080808123342939 , with datetime 2026-09-21T17:01:34.000Z. The original question is Where do you think Grok 4.7 will land? The options are Better, On trend, and Worse. After the vote counts, the page displays “Final results.”
Grok 4.7 has no net-improvement percentage. The chart only shows the label Grok 4.7, with three dashed lines pointing to A Better, B On trend, and C Worse. Its position on the horizontal axis is Coming soon.
The post explicitly says Scores coming soon as you vote on @arena, which means the score has not been published.
The prediction poll has ended, with 412 votes in total: Better 47.1%, On trend 33%, and Worse 19.9%. This is a vote about where it will land, not an Agent Arena net improvement score.
Main-post page snapshot (not a model score): 12 replies, 10 reposts, 203 likes, 13 bookmarks, and a view count displayed as 2.4万.
Historical points already plotted (vertical axis NET IMPROVEMENT (%); these are not Grok 4.7):
| Chart label | Net improvement | Horizontal-axis position |
|---|---|---|
| Grok 4.3 (High) | -8.4% | 8 Jun 2026 |
| Grok Build 0.1 | -6.4% | Drawn between Grok 4.3 (High) and Grok 4.5, with no separate date |
| Grok 4.5 | 5.0% | 13 Jul 2026 |
| Grok 4.6 (xHigh) | 5.3% | 27 Aug 2026 |
| Grok 4.7 | No score plotted | Coming soon |
The three dashed lines for Grok 4.7 have no percentages. Error bars are visible; their values are not stated.
Prediction poll in the same thread (final results):
| Option | Vote share |
|---|---|
| Better | 47.1% |
| On trend | 33% |
| Worse | 19.9% |
The vote total is 412. Engagement on that reply is 0 replies, 0 reposts, 10 likes, and 4,221 views.
The post also quotes https://x.com/arena/status/2102074819743453410 . On this page that quote appears in translated form and is truncated. The visible gist is that Grok 4.7 has entered Agent Arena. The quote card on this page does not give a Grok 4.7 score.
Grok 4.6 (xHigh)'s 5.3%, Grok 4.5's 5.0%, and any negative score on the chart must not be written as Grok 4.7's net improvement.
The 412 votes are a prediction of where 4.7 will land. The page itself states that the score has not been published. Better at 47.1% is not a win rate and is not a leaderboard rank.
The footnote on the historical points says these are results from 7 days after release, but the page gives no task definition, sample size, confidence-interval figures, or reasoning parameters. The error bars cannot be filled in as specific plus-or-minus values.
Grok Build 0.1's -6.4% can be read off the chart. Its date is not labeled separately, so it cannot be assigned to 8 Jun 2026 or 13 Jul 2026.
The main post's view count is shown only as 2.4万; the page does not give a more precise integer.
This is a snapshot of the post and the poll as of 2026-09-22. If Arena later publishes a Grok 4.7 score, that score is outside the scope of this page.
Open https://x.com/arena/status/2102080801462689999 . This collection used a logged-in X session and did not encounter a login wall or a captcha.
Click “Show original” on the main post and check the English body and Scores coming soon.
The image in the post is https://pbs.twimg.com/media/HSwWqqnaAAAPTkt?format=png&name=large . Record only the percentages and dates printed on the chart.
The next reply on the same page, https://x.com/arena/status/2102080808123342939 , is the prediction poll. After clicking “Show original” on that reply, the options are Better, On trend, and Worse.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X (@arena) · Arena.ai (@arena) · Original publication date 2026-09-21 · Site edit date 2026-09-22
Open original sourceGrok 4.7
Download the Tabbit client to check model access