Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Grok 4.7 · Official source · Vendor report

xAI Official Release: Grok 4.7 Benchmark Scores and Capability Positioning

On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.

Official sourceVendor reportEdited 2026-09-22

Test conditions

Source-specific observation
On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
Published conditions
Treating the scores on this page as a third-party rerun; replacing the speed and price slogans in the headline with a measured comparison experiment; folding results for the fast variant, Grok 4.6, or other reasoning tiers into Grok 4.7 xHigh; comparing win rates or cost efficiency without a sample size, a harness。

Key data and applicable tasks

One-sentence takeaway

On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.

Use cases

  • Tasks suitable for a judgment: Confirming how xAI describes Grok 4.7’s product positioning, which benchmark names and scores it published, which reasoning-effort label the comparison table uses, and the listed input and output token prices.

  • Tasks unsuitable for extrapolation: Treating the scores on this page as a third-party rerun; replacing the speed and price slogans in the headline with a measured comparison experiment; folding results for the fast variant, Grok 4.6, or other reasoning tiers into Grok 4.7 xHigh; comparing win rates or cost efficiency without a sample size, a harness, and a scoring script.

  • Applicable model versions: The body is about Grok 4.7. The comparison-table column header is Grok 4.7 xHigh, and the DeepSWE cell is additionally labeled high effort. The scatter plot also shows Grok 4.7 Extra High, High, Medium, and Low. Grok 4.6 appears only in the comparison column and in three bar charts, with the tier labeled High / (high). The page mentions a “fast variant.” The name “Grok 4.7 Fast” does not appear, and the page gives that variant no benchmark scores.

  • Test environment or client: A shared test harness is not stated. The availability section says that, on that day, the model could be used in Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers, and cloud platforms. The training description mentions the Grok Bot harness. The page does not present that harness as the runtime environment for the benchmarks above.

  • Reasoning tiers and parameters: Only the labels visible in each table are recorded. Temperature, top-p, random seed, context length, maximum output, tool configuration, and sample size are all unspecified. The page also does not give an API model ID.

Evaluation method

This is a vendor release note. The page does not provide a protocol that can be rerun independently. This note only transcribes the claims and numbers read from the original page on 2026-09-22.

The only settings that are visible are these:

  • The main comparison table’s column headers name four subjects: Grok 4.7 xHigh, Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max.

  • The caption states that the DeepSWE asterisk on Grok 4.7 denotes a high-effort score. That asterisk is not placed on the other benchmarks in the same column.

  • The CursorBench 4.0 scatter plot can be switched among Cost, Tokens, and Steps. Each point’s accessible name gives the model, the tier, and the score, plus average cost per task, average output tokens, or average step count.

  • The three bar charts below use the tabs Professional knowledge work, Multi-hour office work, and Electrical engineering. A shared caption calls them GDPval, AA Briefcase, and EEBench. The comparison set is Grok 4.7, Grok 4.6, Fable 5.1, and GPT-6 Astra.

  • The safety section names LatchBio’s biosafety benchmark and HackerBench v0.3, which xAI calls “our benchmark.”

Costs, token counts, step counts, and bar-chart scores on the interactive charts change when the page is updated. The readings below are labeled as a 2026-09-22 snapshot. Axis scales (for example 20%–55%, $0–$18, and Elo 0–1500) are scale markings, not model scores.

Sample size, harness name and version, item list, scoring script, confidence interval, and evaluation run date are all unspecified.

Key results

Vendor positioning, and claims that are not paired with a measurement table:

  • The page says Grok 4.7 is its strongest coding and knowledge-work model, that it can work longer on difficult tasks, and that it checks its own output more carefully.

  • The headline and summary say “Twice as fast, at half the price of comparable models.” The comparison set, the speed unit, and the price basis are unspecified.

  • The body says it is offered at the same price and the same speed as Grok 4.6. In the comparison table, both columns list $2 input and $6 output per million tokens. The page does not give a tokens/s figure or a latency measurement.

  • Relative to Grok 4.6, the page says Grok 4.7 uses a new, larger base, that reinforcement learning ran longer, that the task mix is harder and skewed toward multi-hour problems, that self-checking and long context are better, and that it was trained natively for the Grok Bot harness. These sentences do not include a parameter count, a context window, or a training-step count.

  • The page says Grok 4.7 is better at documents and presentations, that it outperforms Grok 4.6 on GDPval and AA Briefcase, and that it is comparable to other frontier models.

  • The safety section says it uses an entirely new safeguard stack and is the strongest, among models xAI itself has tested, on refusal and jailbreak resistance. In dual-use areas such as cybersecurity and biology, it leads on both utility for benign tasks and refusal of dangerous requests. Aside from the two numbers below, no comparative scores are published.

  • The pricing section says there is also a fast variant, with twice the output speed and twice the price. The page does not give that variant’s token list price, a measured speed, or a benchmark score.

Official scores stated directly on the page. All of them lack a sample size and a harness:

  • CursorBench 4.0: In the comparison table, Grok 4.7 xHigh is 46.3%, Grok 4.6 High is 40.4%, GPT-5.6 Sol Max is 41.7%, and Fable 5.1 Max is 51.8%. The scatter plot labels the same group of 46.3% / 41.7% / 51.8% as Extra High / Max / Max, and it shows additional tiers.

  • DeepSWE v1.1: Grok 4.7 is 71.0%, and the page explicitly labels it high effort; Grok 4.6 High is 65.2%; GPT-5.6 Sol Max is 72.7%; Fable 5.1 Max is 70.0%.

  • The main comparison table also includes EEBench, AA Briefcase v1.1, Terminal-Bench 4.0, the Harvey Legal Agent Benchmark, and HealthBench Professional.

  • The bar charts also give GDPval Elo, plus AA Briefcase and EEBench readings that include GPT-6 Astra (max). GPT-6 Astra is not in the main comparison table or the CursorBench scatter plot.

  • LatchBio biosafety benchmark: 62.4%. The page does not explain the numerator or the denominator of that percentage, or the comparison models’ scores.

  • HackerBench v0.3: The page says only 3.3% of risky dual-use prompts were allowed through, and says legitimate security work is rarely blocked. The allow rate for legitimate requests is unspecified.

Raw data

1. Main comparison table

Column order matches the page. The Grok 4.7 column header is xHigh. Only DeepSWE is recorded as high effort, following the caption. The other cells keep the column-header tier. Sample size, harness, and evaluation date are all unspecified.

ItemGrok 4.7Grok 4.6GPT-5.6 SolFable 5.1
Reasoning tier (column header)xHigh; DeepSWE is high effortHighMaxMax
Input price (per million tokens)$2$2$4$10
Output price (per million tokens)$6$6$20$50
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0% (high effort)65.2%72.7%70.0%
EEBench64.0%53.0%39.4%56.4%
AA Briefcase v1.11,6571,5461,4871,678
Terminal-Bench 4.038.0%20.3%37.3%57.9%
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
HealthBench Professional56.7%48.5%60.5%62.1%

The category labels the page attaches to these rows are Software engineering, Electrical engineering, Multi-hour office work, Multi-hour terminal work, Legal work, and Clinical reasoning. AA Briefcase v1.1 and the bar chart later in this note use Elo-style large numbers. The page does not explain the unit again beside the comparison table.

The pricing section also writes “priced starting at” $2 input and $6 output per million tokens. The fast variant does not have its own price row.

2. CursorBench 4.0 scatter plot (2026-09-22 snapshot)

The chart’s accessible name is “CursorBench 4.0 score by average cost / output tokens / steps per task.” The models visible on the points are Fable 5.1, Opus 5, Grok 4.7, GPT-5.6 Sol, and Sonnet 5. Grok 4.6 is not on this chart. Grok 4.7 has no Max point. For every point, sample size, harness, input tokens, and evaluation date are unspecified. Average cost is cost per task. That is not the same basis as the per-million-token list price in the previous section.

Grok 4.7 Extra High’s 46.3% is the same CursorBench 4.0 score as Grok 4.7 xHigh in the comparison table. The page uses both spellings, Extra High and xHigh, and does not give a separate API parameter name.

Model and on-page tierScoreAverage cost per taskAverage output tokensAverage steps
Grok 4.7 Extra High46.3%$6.0170,14188
Grok 4.7 High43.9%$4.6956,38271
Grok 4.7 Medium41.6%$3.4936,68360
Grok 4.7 Low33.1%$1.5815,67740
Fable 5.1 Max51.8%$17.28117,236128
Fable 5.1 Extra High51.6%$13.0187,294101
Fable 5.1 High49.2%$9.0858,43877
Fable 5.1 Medium46.8%$7.0545,41163
Fable 5.1 Low45.1%$5.4434,79551
Opus 5 Max46.6%$11.9585,384106
Opus 5 Extra High46.1%$11.4380,094103
Opus 5 High44.7%$9.0061,40586
Opus 5 Medium43.3%$6.9445,27272
Opus 5 Low40.7%$4.8731,99557
GPT-5.6 Sol Max41.7%$8.2342,94499
GPT-5.6 Sol Extra High37.7%$4.4024,72955
GPT-5.6 Sol High35.7%$2.8516,17441
GPT-5.6 Sol Medium31.1%$1.7710,11132
GPT-5.6 Sol Low24.6%$0.874,88521
Sonnet 5 Max34.1%$7.17149,257140
Sonnet 5 Extra High32.0%$4.5583,373102
Sonnet 5 High30.8%$3.4861,14685
Sonnet 5 Medium28.0%$2.3139,11465
Sonnet 5 Low24.1%$1.3923,77246

3. GDPval, AA Briefcase, and EEBench bar charts (2026-09-22 snapshot)

The tier notation on the three charts is (xhigh), (high), and (max). Sample size, harness, and evaluation date are unspecified. The bar charts use GPT-6 Astra (max). The main comparison table uses GPT-5.6 Sol Max. Do not merge the two comparison sets into one row.

ChartMetricGrok 4.7 (xhigh)Grok 4.6 (high)Fable 5.1 (max)GPT-6 Astra (max)
Professional knowledge work / GDPvalElo1,6951,6051,7351,542
Multi-hour office work / AA BriefcaseElo1,6571,5461,6781,569
Electrical engineering / EEBenchAccuracy64.0%53.0%56.4%69.3%

The office chart’s accessible name is “Multi-hour office work scores by model.” The label itself does not also print “AA Briefcase v1.1.” The values 1,657, 1,546, and 1,678 match the corresponding AA Briefcase v1.1 cells in the main comparison table. The main comparison table also has 1,487 for GPT-5.6 Sol Max. The bar chart does not include that point.

The electrical-engineering chart’s unit is Accuracy. The values 64.0%, 53.0%, and 56.4% match the corresponding EEBench cells in the main comparison table. The main comparison table also has 39.4% for GPT-5.6 Sol Max. The 69.3% on the bar chart belongs to GPT-6 Astra (max), not to GPT-5.6 Sol.

GDPval appears only on the bar chart. It is not in the main comparison table.

4. Safety numbers

ItemValue given on the pageUnspecified fields
LatchBio biosafety benchmarkGrok 4.7 is 62.4%. The page says it leads on this benchmarkModel-version tier, comparison-model scores, sample size, harness, metric definition, evaluation date
HackerBench v0.33.3% of risky dual-use prompts were allowed through. The page says this is the benchmark xAI uses for risky and malicious cyber tasks, and says its safety is the highestReasoning tier, total prompt count, block rate for legitimate security work, comparison-model scores, harness, evaluation date
Refusal and jailbreak resistanceThe page says it is the strongest model it has testedBenchmark name, score, sample size
Red-team capabilityThe page says it has given invite-only access to some cybersecurity partners for defensive researchNo public score is given

Conclusions and limitations

Every benchmark number on this page is an official result in xAI’s release post. Independent measurement, third-party rerun logs, and statistical significance are all unspecified. This release post cannot be written up as an independent rerun of Grok 4.7.

Record the versions separately:

  • Grok 4.7’s main comparison column is xHigh, but the 71.0% on DeepSWE v1.1 is labeled high effort on its own. The scatter plot also publishes CursorBench results for the four tiers Extra High, High, Medium, and Low. 46.3% cannot stand in for every tier.

  • “Grok 4.7 Fast” does not appear on the page. The fast variant has only the sentence about twice the output speed and twice the price. It has no benchmark score and no separate list price. The Grok 4.7 scores on the scatter plot and the comparison table do not apply to this variant.

  • Grok 4.6’s published tier is High / (high). Its CursorBench 4.0 score is 40.4%, which is not the same number as any Grok 4.7 tier.

There is more than one comparison set. Between them, the main comparison table and the CursorBench scatter plot bring in GPT-5.6 Sol, Opus 5, and Sonnet 5. The three bar charts bring in GPT-6 Astra and drop GPT-5.6 Sol. On EEBench, GPT-5.6 Sol Max at 39.4% and GPT-6 Astra (max) at 69.3% are therefore both visible, and they come from different comparison sets.

Keep price and speed separate, following the original sentences:

  • The comparison table supports the reading that Grok 4.7 and Grok 4.6 share the same list price of $2 / $6.

  • “The same speed as Grok 4.6,” “twice as fast as comparable models, at half the price,” and the fast variant’s “twice the speed, twice the price” are not paired with a speed table or a named comparison set.

  • Average cost per task on the scatter plot is CursorBench task spend. It cannot be converted into the per-million-token list price, and it does not represent the fast variant.

Safety conclusions stay with the page’s original sentences. 62.4% and 3.3% have no sample size. HackerBench is written as xAI’s own benchmark. Invite-only red-team access is not a public score.

The dynamic charts are a page snapshot from 2026-09-22. The article’s publication date is 2026-09-21. If the page later updates the scatter plot or the bar charts, the reopened original is the source of record.

Reproduction notes

These scores cannot be reproduced from this page alone. The page does not state the item-set version, sample size, full prompts, harness, tool schema, temperature, random seed, scoring script, failed samples, or confidence interval, and it does not give an API model ID.

If a separate rerun is done later, the model version and reasoning tier in use at that time need to be fixed first. Grok 4.7 xHigh, high effort, Extra High, High, Medium, Low, and the fast variant should be separate runs. Grok 4.6 should be listed on its own. This file does not provide the results of such a rerun.

What this supports

  • Confirming how xAI describes Grok 4.7’s product positioning, which benchmark names and scores it published, which reasoning-effort label the comparison table uses, and the listed input and output token prices.

What this does not support

  • Treating the scores on this page as a third-party rerun; replacing the speed and price slogans in the headline with a measured comparison experiment; folding results for the fast variant, Grok 4.6, or other reasoning tiers into Grok 4.7 xHigh; comparing win rates or cost efficiency without a sample size, a harness。

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

xAI News (the browser title suffix is SpaceXAI; the footer copyright is SpaceXAI LLC) · xAI (the organization author in schema.org; the body has no individual byline) · Original publication date 2026-09-21 · Site edit date 2026-09-22

Open original source

Grok 4.7

Compare Grok 4.7 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Full review · English

Grok 4.7 Review: Same $2/$6 Price, About Twice the Tokens

A Grok 4.7 review of the unchanged $2/$6 rates, the jump to about 81k output tokens, and which workloads justify the extra work.

Pricing · English

Grok 4.7 Pricing: The $2/$6 Card and the Real Bill

Grok 4.7 keeps Grok 4.6's $2, $0.50, and $6 API rates. Effort, the 200k cliff, Fast, and Cursor's 256k line decide the bill.

Comparison · English

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

Related reviews

xAI Official Model Card: Grok 4.7 Safety Evaluation and Use BoundariesThe official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.X: Artificial Analysis’s AA-Briefcase Chart — Grok 4.7 (xhigh) Composite Elo 1657On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).Arena: Grok 4.7 Has No Score Yet on the Agent Arena Net Improvement ChartAs of 2026-09-22, this Arena post has not published an Agent Arena net improvement score for Grok 4.7. The chart labels Grok 4.7 as Coming soon, the post says Scores coming soon, and the poll in the same thread is only a prediction by 412 people about where it will land.Same-Task Cursor Ultra Usage for Grok 4.7 and Grok 4.6On Cursor Ultra, at extra high, with fast mode off, the author estimates from how fast the same task reduced plan usage that Grok 4.7 consumes about 2.5 times as much as Grok 4.6. This is a short timing on a single account and cannot be treated as a quality benchmark.Grok 4.7 API setup on OpenRouterThe OpenRouter model page labels x-ai/grok-4.7 as SpaceXAI's Grok 4.7, lists input / output prices of $1.60 / $4.80 per million tokens, and gives OpenRouter SDK and cURL examples; the reasoning-level field in the request body, and the list prices for Low, Medium, and High, were not read in this collection.xAI Official Documentation: Grok 4.7 API Parameters and Reasoning LevelsThe model name on the public xAI API is grok-4.7. The reasoning levels listed in the documentation are Low, medium, high (default), or xhigh, with an input price of $2.00 / 1M tokens and an output price of $6.00 / 1M tokens. Grok 4.7 Fast is written as a faster deployment of the same model, billed at twice the standard token price, and it appears only in Cursor and Grok Build.Box's prompt and result for reviewing the Merewick claim with Grok 4.7 in AI StudioThis Box post is a 41-second Box Agent preview. In Box AI Studio, Grok 4.7 is selected and a claim-review prompt is entered for the folder "Active Commercial Property Claims", producing the file Claim Reconciliation Review Merewick Coastal Foods ASC-26-0184.md.