Grok 4.7 · Official source · Vendor report
On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
On the release page, xAI positions Grok 4.7 as a coding and knowledge-work model and self-reports scores on CursorBench 4.0, DeepSWE v1.1, EEBench, and other benchmarks. These figures are vendor claims visible when the original page was opened on 2026-09-22. This note does not include an independent rerun.
Tasks suitable for a judgment: Confirming how xAI describes Grok 4.7’s product positioning, which benchmark names and scores it published, which reasoning-effort label the comparison table uses, and the listed input and output token prices.
Tasks unsuitable for extrapolation: Treating the scores on this page as a third-party rerun; replacing the speed and price slogans in the headline with a measured comparison experiment; folding results for the fast variant, Grok 4.6, or other reasoning tiers into Grok 4.7 xHigh; comparing win rates or cost efficiency without a sample size, a harness, and a scoring script.
Applicable model versions: The body is about Grok 4.7. The comparison-table column header is Grok 4.7 xHigh, and the DeepSWE cell is additionally labeled high effort. The scatter plot also shows Grok 4.7 Extra High, High, Medium, and Low. Grok 4.6 appears only in the comparison column and in three bar charts, with the tier labeled High / (high). The page mentions a “fast variant.” The name “Grok 4.7 Fast” does not appear, and the page gives that variant no benchmark scores.
Test environment or client: A shared test harness is not stated. The availability section says that, on that day, the model could be used in Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers, and cloud platforms. The training description mentions the Grok Bot harness. The page does not present that harness as the runtime environment for the benchmarks above.
Reasoning tiers and parameters: Only the labels visible in each table are recorded. Temperature, top-p, random seed, context length, maximum output, tool configuration, and sample size are all unspecified. The page also does not give an API model ID.
This is a vendor release note. The page does not provide a protocol that can be rerun independently. This note only transcribes the claims and numbers read from the original page on 2026-09-22.
The only settings that are visible are these:
The main comparison table’s column headers name four subjects: Grok 4.7 xHigh, Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max.
The caption states that the DeepSWE asterisk on Grok 4.7 denotes a high-effort score. That asterisk is not placed on the other benchmarks in the same column.
The CursorBench 4.0 scatter plot can be switched among Cost, Tokens, and Steps. Each point’s accessible name gives the model, the tier, and the score, plus average cost per task, average output tokens, or average step count.
The three bar charts below use the tabs Professional knowledge work, Multi-hour office work, and Electrical engineering. A shared caption calls them GDPval, AA Briefcase, and EEBench. The comparison set is Grok 4.7, Grok 4.6, Fable 5.1, and GPT-6 Astra.
The safety section names LatchBio’s biosafety benchmark and HackerBench v0.3, which xAI calls “our benchmark.”
Costs, token counts, step counts, and bar-chart scores on the interactive charts change when the page is updated. The readings below are labeled as a 2026-09-22 snapshot. Axis scales (for example 20%–55%, $0–$18, and Elo 0–1500) are scale markings, not model scores.
Sample size, harness name and version, item list, scoring script, confidence interval, and evaluation run date are all unspecified.
Vendor positioning, and claims that are not paired with a measurement table:
The page says Grok 4.7 is its strongest coding and knowledge-work model, that it can work longer on difficult tasks, and that it checks its own output more carefully.
The headline and summary say “Twice as fast, at half the price of comparable models.” The comparison set, the speed unit, and the price basis are unspecified.
The body says it is offered at the same price and the same speed as Grok 4.6. In the comparison table, both columns list $2 input and $6 output per million tokens. The page does not give a tokens/s figure or a latency measurement.
Relative to Grok 4.6, the page says Grok 4.7 uses a new, larger base, that reinforcement learning ran longer, that the task mix is harder and skewed toward multi-hour problems, that self-checking and long context are better, and that it was trained natively for the Grok Bot harness. These sentences do not include a parameter count, a context window, or a training-step count.
The page says Grok 4.7 is better at documents and presentations, that it outperforms Grok 4.6 on GDPval and AA Briefcase, and that it is comparable to other frontier models.
The safety section says it uses an entirely new safeguard stack and is the strongest, among models xAI itself has tested, on refusal and jailbreak resistance. In dual-use areas such as cybersecurity and biology, it leads on both utility for benign tasks and refusal of dangerous requests. Aside from the two numbers below, no comparative scores are published.
The pricing section says there is also a fast variant, with twice the output speed and twice the price. The page does not give that variant’s token list price, a measured speed, or a benchmark score.
Official scores stated directly on the page. All of them lack a sample size and a harness:
CursorBench 4.0: In the comparison table, Grok 4.7 xHigh is 46.3%, Grok 4.6 High is 40.4%, GPT-5.6 Sol Max is 41.7%, and Fable 5.1 Max is 51.8%. The scatter plot labels the same group of 46.3% / 41.7% / 51.8% as Extra High / Max / Max, and it shows additional tiers.
DeepSWE v1.1: Grok 4.7 is 71.0%, and the page explicitly labels it high effort; Grok 4.6 High is 65.2%; GPT-5.6 Sol Max is 72.7%; Fable 5.1 Max is 70.0%.
The main comparison table also includes EEBench, AA Briefcase v1.1, Terminal-Bench 4.0, the Harvey Legal Agent Benchmark, and HealthBench Professional.
The bar charts also give GDPval Elo, plus AA Briefcase and EEBench readings that include GPT-6 Astra (max). GPT-6 Astra is not in the main comparison table or the CursorBench scatter plot.
LatchBio biosafety benchmark: 62.4%. The page does not explain the numerator or the denominator of that percentage, or the comparison models’ scores.
HackerBench v0.3: The page says only 3.3% of risky dual-use prompts were allowed through, and says legitimate security work is rarely blocked. The allow rate for legitimate requests is unspecified.
Column order matches the page. The Grok 4.7 column header is xHigh. Only DeepSWE is recorded as high effort, following the caption. The other cells keep the column-header tier. Sample size, harness, and evaluation date are all unspecified.
| Item | Grok 4.7 | Grok 4.6 | GPT-5.6 Sol | Fable 5.1 |
|---|---|---|---|---|
| Reasoning tier (column header) | xHigh; DeepSWE is high effort | High | Max | Max |
| Input price (per million tokens) | $2 | $2 | $4 | $10 |
| Output price (per million tokens) | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0% (high effort) | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
The category labels the page attaches to these rows are Software engineering, Electrical engineering, Multi-hour office work, Multi-hour terminal work, Legal work, and Clinical reasoning. AA Briefcase v1.1 and the bar chart later in this note use Elo-style large numbers. The page does not explain the unit again beside the comparison table.
The pricing section also writes “priced starting at” $2 input and $6 output per million tokens. The fast variant does not have its own price row.
The chart’s accessible name is “CursorBench 4.0 score by average cost / output tokens / steps per task.” The models visible on the points are Fable 5.1, Opus 5, Grok 4.7, GPT-5.6 Sol, and Sonnet 5. Grok 4.6 is not on this chart. Grok 4.7 has no Max point. For every point, sample size, harness, input tokens, and evaluation date are unspecified. Average cost is cost per task. That is not the same basis as the per-million-token list price in the previous section.
Grok 4.7 Extra High’s 46.3% is the same CursorBench 4.0 score as Grok 4.7 xHigh in the comparison table. The page uses both spellings, Extra High and xHigh, and does not give a separate API parameter name.
| Model and on-page tier | Score | Average cost per task | Average output tokens | Average steps |
|---|---|---|---|---|
| Grok 4.7 Extra High | 46.3% | $6.01 | 70,141 | 88 |
| Grok 4.7 High | 43.9% | $4.69 | 56,382 | 71 |
| Grok 4.7 Medium | 41.6% | $3.49 | 36,683 | 60 |
| Grok 4.7 Low | 33.1% | $1.58 | 15,677 | 40 |
| Fable 5.1 Max | 51.8% | $17.28 | 117,236 | 128 |
| Fable 5.1 Extra High | 51.6% | $13.01 | 87,294 | 101 |
| Fable 5.1 High | 49.2% | $9.08 | 58,438 | 77 |
| Fable 5.1 Medium | 46.8% | $7.05 | 45,411 | 63 |
| Fable 5.1 Low | 45.1% | $5.44 | 34,795 | 51 |
| Opus 5 Max | 46.6% | $11.95 | 85,384 | 106 |
| Opus 5 Extra High | 46.1% | $11.43 | 80,094 | 103 |
| Opus 5 High | 44.7% | $9.00 | 61,405 | 86 |
| Opus 5 Medium | 43.3% | $6.94 | 45,272 | 72 |
| Opus 5 Low | 40.7% | $4.87 | 31,995 | 57 |
| GPT-5.6 Sol Max | 41.7% | $8.23 | 42,944 | 99 |
| GPT-5.6 Sol Extra High | 37.7% | $4.40 | 24,729 | 55 |
| GPT-5.6 Sol High | 35.7% | $2.85 | 16,174 | 41 |
| GPT-5.6 Sol Medium | 31.1% | $1.77 | 10,111 | 32 |
| GPT-5.6 Sol Low | 24.6% | $0.87 | 4,885 | 21 |
| Sonnet 5 Max | 34.1% | $7.17 | 149,257 | 140 |
| Sonnet 5 Extra High | 32.0% | $4.55 | 83,373 | 102 |
| Sonnet 5 High | 30.8% | $3.48 | 61,146 | 85 |
| Sonnet 5 Medium | 28.0% | $2.31 | 39,114 | 65 |
| Sonnet 5 Low | 24.1% | $1.39 | 23,772 | 46 |
The tier notation on the three charts is (xhigh), (high), and (max). Sample size, harness, and evaluation date are unspecified. The bar charts use GPT-6 Astra (max). The main comparison table uses GPT-5.6 Sol Max. Do not merge the two comparison sets into one row.
| Chart | Metric | Grok 4.7 (xhigh) | Grok 4.6 (high) | Fable 5.1 (max) | GPT-6 Astra (max) |
|---|---|---|---|---|---|
| Professional knowledge work / GDPval | Elo | 1,695 | 1,605 | 1,735 | 1,542 |
| Multi-hour office work / AA Briefcase | Elo | 1,657 | 1,546 | 1,678 | 1,569 |
| Electrical engineering / EEBench | Accuracy | 64.0% | 53.0% | 56.4% | 69.3% |
The office chart’s accessible name is “Multi-hour office work scores by model.” The label itself does not also print “AA Briefcase v1.1.” The values 1,657, 1,546, and 1,678 match the corresponding AA Briefcase v1.1 cells in the main comparison table. The main comparison table also has 1,487 for GPT-5.6 Sol Max. The bar chart does not include that point.
The electrical-engineering chart’s unit is Accuracy. The values 64.0%, 53.0%, and 56.4% match the corresponding EEBench cells in the main comparison table. The main comparison table also has 39.4% for GPT-5.6 Sol Max. The 69.3% on the bar chart belongs to GPT-6 Astra (max), not to GPT-5.6 Sol.
GDPval appears only on the bar chart. It is not in the main comparison table.
| Item | Value given on the page | Unspecified fields |
|---|---|---|
| LatchBio biosafety benchmark | Grok 4.7 is 62.4%. The page says it leads on this benchmark | Model-version tier, comparison-model scores, sample size, harness, metric definition, evaluation date |
| HackerBench v0.3 | 3.3% of risky dual-use prompts were allowed through. The page says this is the benchmark xAI uses for risky and malicious cyber tasks, and says its safety is the highest | Reasoning tier, total prompt count, block rate for legitimate security work, comparison-model scores, harness, evaluation date |
| Refusal and jailbreak resistance | The page says it is the strongest model it has tested | Benchmark name, score, sample size |
| Red-team capability | The page says it has given invite-only access to some cybersecurity partners for defensive research | No public score is given |
Every benchmark number on this page is an official result in xAI’s release post. Independent measurement, third-party rerun logs, and statistical significance are all unspecified. This release post cannot be written up as an independent rerun of Grok 4.7.
Record the versions separately:
Grok 4.7’s main comparison column is xHigh, but the 71.0% on DeepSWE v1.1 is labeled high effort on its own. The scatter plot also publishes CursorBench results for the four tiers Extra High, High, Medium, and Low. 46.3% cannot stand in for every tier.
“Grok 4.7 Fast” does not appear on the page. The fast variant has only the sentence about twice the output speed and twice the price. It has no benchmark score and no separate list price. The Grok 4.7 scores on the scatter plot and the comparison table do not apply to this variant.
Grok 4.6’s published tier is High / (high). Its CursorBench 4.0 score is 40.4%, which is not the same number as any Grok 4.7 tier.
There is more than one comparison set. Between them, the main comparison table and the CursorBench scatter plot bring in GPT-5.6 Sol, Opus 5, and Sonnet 5. The three bar charts bring in GPT-6 Astra and drop GPT-5.6 Sol. On EEBench, GPT-5.6 Sol Max at 39.4% and GPT-6 Astra (max) at 69.3% are therefore both visible, and they come from different comparison sets.
Keep price and speed separate, following the original sentences:
The comparison table supports the reading that Grok 4.7 and Grok 4.6 share the same list price of $2 / $6.
“The same speed as Grok 4.6,” “twice as fast as comparable models, at half the price,” and the fast variant’s “twice the speed, twice the price” are not paired with a speed table or a named comparison set.
Average cost per task on the scatter plot is CursorBench task spend. It cannot be converted into the per-million-token list price, and it does not represent the fast variant.
Safety conclusions stay with the page’s original sentences. 62.4% and 3.3% have no sample size. HackerBench is written as xAI’s own benchmark. Invite-only red-team access is not a public score.
The dynamic charts are a page snapshot from 2026-09-22. The article’s publication date is 2026-09-21. If the page later updates the scatter plot or the bar charts, the reopened original is the source of record.
These scores cannot be reproduced from this page alone. The page does not state the item-set version, sample size, full prompts, harness, tool schema, temperature, random seed, scoring script, failed samples, or confidence interval, and it does not give an API model ID.
If a separate rerun is done later, the model version and reasoning tier in use at that time need to be fixed first. Grok 4.7 xHigh, high effort, Extra High, High, Medium, Low, and the fast variant should be separate runs. Grok 4.6 should be listed on its own. This file does not provide the results of such a rerun.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
xAI News (the browser title suffix is SpaceXAI; the footer copyright is SpaceXAI LLC) · xAI (the organization author in schema.org; the body has no individual byline) · Original publication date 2026-09-21 · Site edit date 2026-09-22
Open original sourceGrok 4.7
Download the Tabbit client to check model access