Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Grok 4.7 · Community source · Platform telemetry

Arena: Grok 4.7 Has No Score Yet on the Agent Arena Net Improvement Chart

As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.

Community sourcePlatform telemetryEdited 2026-09-22

Test conditions

Source-specific observation
As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.
Published conditions
It cannot be used to judge Grok 4.7's actual net improvement, ranking, or blind-evaluation win rate, and it cannot be extended to quality on text, vision, code, or document tasks.

Key data and applicable tasks

One-sentence takeaway

As of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.

Use cases

  • Tasks suitable for a judgment: Whether Arena has already published a net improvement score for Grok 4.7, and whether, at collection time, the community predicted 4.7 relative to the historical trend as Better, On trend, or Worse.

  • Tasks unsuitable for extrapolation: It cannot be used to judge Grok 4.7's actual net improvement, ranking, or blind-evaluation win rate, and it cannot be extended to quality on text, vision, code, or document tasks.

  • Applicable model version: The model discussed in the post that still has no score is Grok 4.7. The points already plotted belong to Grok 4.3 (High), Grok Build 0.1, Grok 4.5, and Grok 4.6 (xHigh), and must not be written as Grok 4.7 results.

  • Test environment or client: The chart title is Agent Arena, the source credit reads SOURCE: AGENT ARENA, and the footnote is RESULTS 7 DAYS AFTER RELEASE. Grok 4.7 has no corresponding result point on this chart.

  • Reasoning tiers and parameters: Not stated for Grok 4.7. The chart labels only the older versions Grok 4.3 (High) and Grok 4.6 (xHigh). Temperature, tool configuration, and sample size are not stated.

Evaluation method

This is not a Grok 4.7 leaderboard row with a public vote count, and it is not a controlled evaluation of Grok 4.7. In the post, @arena shows a historical net-improvement line chart and starts a prediction poll in a reply on the same page. The original post text is: Grok 4.7 by @SpaceXAI just dropped. Looking at how past Grok versions have trended on Agent Arena's net improvement score, where do you think 4.7 lands? Poll below. Scores coming soon as you vote on @arena!

The vertical axis is NET IMPROVEMENT (%), with visible ticks at 10%, 5%, 0%, -5%, and -10%. The horizontal axis is labeled 8 Jun 2026, 13 Jul 2026, 27 Aug 2026, and Coming soon. The plotted points have vertical error bars, but the chart does not print values for those bars. The footnote reads RESULTS 7 DAYS AFTER RELEASE. It does not state the task set, the prompts, the sample size, or the vote counts that produced these historical scores.

The prediction poll is the next @arena reply on the same page: https://x.com/arena/status/2102080808123342939 , with datetime 2026-09-21T17:01:34.000Z. The original question is Where do you think Grok 4.7 will land? The options are Better, On trend, and Worse. After the vote counts, the page displays “Final results.”

Key results

  • Grok 4.7 has no net-improvement percentage. The chart only shows the label Grok 4.7, with three dashed lines pointing to A Better, B On trend, and C Worse. Its position on the horizontal axis is Coming soon.

  • The post explicitly says Scores coming soon as you vote on @arena, which means the score has not been published.

  • The prediction poll has ended, with 412 votes in total: Better 47.1%, On trend 33%, and Worse 19.9%. This is a vote about where it will land, not an Agent Arena net improvement score.

Raw data

Main-post page snapshot (not a model score): 12 replies, 10 reposts, 203 likes, 13 bookmarks, and a view count displayed as 2.4万.

Historical points already plotted (vertical axis NET IMPROVEMENT (%); these are not Grok 4.7):

Chart labelNet improvementHorizontal-axis position
Grok 4.3 (High)-8.4%8 Jun 2026
Grok Build 0.1-6.4%Drawn between Grok 4.3 (High) and Grok 4.5, with no separate date
Grok 4.55.0%13 Jul 2026
Grok 4.6 (xHigh)5.3%27 Aug 2026
Grok 4.7No score plottedComing soon

The three dashed lines for Grok 4.7 have no percentages. Error bars are visible; their values are not stated.

Prediction poll in the same thread (final results):

OptionVote share
Better47.1%
On trend33%
Worse19.9%

The vote total is 412. Engagement on that reply is 0 replies, 0 reposts, 10 likes, and 4,221 views.

The post also quotes https://x.com/arena/status/2102074819743453410 . On this page that quote appears in translated form and is truncated. The visible gist is that Grok 4.7 has entered Agent Arena. The quote card on this page does not give a Grok 4.7 score.

Conclusions and limitations

  • Grok 4.6 (xHigh)'s 5.3%, Grok 4.5's 5.0%, and any negative score on the chart must not be written as Grok 4.7's net improvement.

  • The 412 votes are a prediction of where 4.7 will land. The page itself states that the score has not been published. Better at 47.1% is not a win rate and is not a leaderboard rank.

  • The footnote on the historical points says these are results from 7 days after release, but the page gives no task definition, sample size, confidence-interval figures, or reasoning parameters. The error bars cannot be filled in as specific plus-or-minus values.

  • Grok Build 0.1's -6.4% can be read off the chart. Its date is not labeled separately, so it cannot be assigned to 8 Jun 2026 or 13 Jul 2026.

  • The main post's view count is shown only as 2.4万; the page does not give a more precise integer.

  • This is a snapshot of the post and the poll as of 2026-09-22. If Arena later publishes a Grok 4.7 score, that score is outside the scope of this page.

Reproduction notes

  1. Open https://x.com/arena/status/2102080801462689999 . This collection used a logged-in X session and did not encounter a login wall or a captcha.

  2. Click “Show original” on the main post and check the English body and Scores coming soon.

  3. The image in the post is https://pbs.twimg.com/media/HSwWqqnaAAAPTkt?format=png&name=large . Record only the percentages and dates printed on the chart.

  4. The next reply on the same page, https://x.com/arena/status/2102080808123342939 , is the prediction poll. After clicking “Show original” on that reply, the options are Better, On trend, and Worse.

What this supports

  • Whether Arena has already published a net improvement score for Grok 4.7, and whether, at collection time, the community predicted 4.7 relative to the historical trend as Better, On trend, or Worse.

What this does not support

  • It cannot be used to judge Grok 4.7's actual net improvement, ranking, or blind-evaluation win rate, and it cannot be extended to quality on text, vision, code, or document tasks.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X (@arena) · Arena.ai (@arena) · Original publication date 2026-09-21 · Site edit date 2026-09-22

Open original source

Grok 4.7

Compare Grok 4.7 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Full review · English

Grok 4.7 Review: Same $2/$6 Price, About Twice the Tokens

A Grok 4.7 review of the unchanged $2/$6 rates, the jump to about 81k output tokens, and which workloads justify the extra work.

Pricing · English

Grok 4.7 Pricing: The $2/$6 Card and the Real Bill

Grok 4.7 keeps Grok 4.6's $2, $0.50, and $6 API rates. Effort, the 200k cliff, Fast, and Cursor's 256k line decide the bill.

Comparison · English

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

Related reviews

X: Artificial Analysis’s AA-Briefcase Chart — Grok 4.7 (xhigh) Composite Elo 1657On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).xAI Official Release: Grok 4.7 Benchmark Scores and Capability PositioningOn the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.xAI Official Model Card: Grok 4.7 Safety Evaluation and Use BoundariesThe official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.Artificial Analysis: Grok 4.7 Intelligence Index and Coding AgentArtificial Analysis measured Grok 4.7 at xhigh and recorded an Intelligence Index v4.3 of 46, and a Coding Agent Index v1.5 of 56 on Grok Build. Knowledge-work Elo is higher than Grok 4.6 (high), and output tokens per Intelligence Index question are about 81k.Box's prompt and result for reviewing the Merewick claim with Grok 4.7 in AI StudioThis Box post is a 41-second Box Agent preview. In Box AI Studio, Grok 4.7 is selected and a claim-review prompt is entered for the folder "Active Commercial Property Claims", producing the file Claim Reconciliation Review Merewick Coastal Foods ASC-26-0184.md.Grok 4.7 API setup on OpenRouterThe OpenRouter model page labels x-ai/grok-4.7 as SpaceXAI's Grok 4.7, lists input / output prices of $1.60 / $4.80 per million tokens, and gives OpenRouter SDK and cURL examples; the reasoning-level field in the request body, and the list prices for Low, Medium, and High, were not read in this collection.xAI Official Documentation: Grok 4.7 API Parameters and Reasoning LevelsThe model name on the public xAI API is grok-4.7. The reasoning levels listed in the documentation are Low, medium, high (default), or xhigh, with an input price of $2.00 / 1M tokens and an output price of $6.00 / 1M tokens. Grok 4.7 Fast is written as a faster deployment of the same model, billed at twice the standard token price, and it appears only in Cursor and Grok Build.