Grok 4.7 · Media / benchmark · Vendor report
The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
The official model card describes Grok 4.7 as a deployed checkpoint after Grok 4.6, aimed at coding, engineering, and office tasks. It lists available channels, the training cutoff date, and capability and safety scores at the xhigh or high tier. The card does not mention a fast variant or an API model ID.
Tasks suitable for a judgment: Confirming how the official card limits use, which channels it lists as available, what training details it provides, and at which reasoning tier the card prints scores for Grok 4.7, Grok 4.6, and comparison models in each table.
Tasks unsuitable for extrapolation: Treating official scores as a third-party rerun; applying a score from one tier to another; counting results for Grok 4.6, Grok 4.5, or comparison models as Grok 4.7 results; using unprotected capability probes as a substitute for post-deployment refusal behavior; or adding fast variants, consumer products, or scatter-plot coordinates not shown in the card to the conclusions.
Applicable model versions: The subject is the final deployed Grok 4.7 checkpoint. The card uses xhigh or high by section, and prints both tiers for CVE-Bench and HackerBench. Comparison columns also include Grok 4.6 (mostly high, with xhigh additionally shown for EEBench and FrontierSWE) and Grok 4.5 (high). The card does not mention “Grok 4.7 Fast” or a “fast variant.”
Test environment or client: No unified test client is specified. The harness named for each item appears in the evaluation method. Availability channels and evaluation harnesses are not the same thing.
Reasoning tiers and parameters: Only the printed xhigh and high tiers, plus max and xhigh for comparison models, are recorded. Temperature, top-p, random seed, context length, maximum output, tool schema, and API model ID are all unspecified.
This is a vendor model card, not an independently rerunnable protocol. The card states that, unless otherwise specified, evaluation results come from the final deployed Grok 4.7 checkpoint.
Capability and safety items are separated. Cybersecurity capability probes run with production safeguards disabled, while refusal is measured separately in HackerBench. The biology and chemistry benchmarks in Section 7 default to production safeguards being disabled so that dual-use capability can be measured. Jailbreaks, general output safety, child safety, CBRN / weapons refusal, self-harm, MASK, and sycophancy measure safeguards or behavior, not unprotected capability.
The card names these runtime conditions:
CursorBench 4.0 gives Grok 4.7's xhigh and high scores only in the body. Its two scatter plots have axes and legends but no per-point numbers.
DeepSWE v1.1: 113 long-horizon tasks, 91 repositories, 5 languages; mini-SWE-agent; Pass@1; evaluated by Datacurve.
Terminal-Bench 4.0: 66 terminal workflows, with a maximum of 8 hours per task; Grok 4.7 uses the Grok Build harness; evaluated by Harbor.
FrontierSWE V2: 34 tasks, a maximum of 20 hours per task, 5 trials per item, reporting mean@5. The public protocol uses Proximal's Proximus harness at maximum effort. Grok's score was evaluated by Proximal Labs.
SWE-Marathon v1.1: 20 tasks, 8 trials per item; each model uses its native agent harness; evaluated by Abundant AI.
Legal Agent Benchmark: Vals AI's 120-question holdout set, Valkyrie harness, internet disabled; Harvey's final score is the average task pass rate from two reviewers.
EEBench: Grok 4.7 uses Grok Build, while comparison models use their respective provider harnesses; evaluated by Atopile; a footnote says this is the benchmark's V1 core corpus.
CADGenBench: reports the generation split; Grok 4.7 uses Grok Build; evaluated by Mecado.
HealthBench Professional: questions rejected by the safety filter score 0. Grok 4.7's result uses Grok 4.6 (high) as the scoring model; the Fable and Astra figures come from the vendors' own model cards.
LatchBio Capabilities v1.0: equal-weight average of 11 capability benchmarks. The card calls this a live suite that is updated over time.
CyberGym and CVE-Bench: no standard safeguards; Grok 4.7 uses Grok Build. The unprotected results for other CyberGym models come from their respective system cards.
HackerBench v0.3: Grok 4.7 uses standard safeguards tracked with the release; other models use the standard safeguards provided by their vendors.
CathedralBench: unrestricted configuration, reporting only the hard subset.
BioSecBench: Grok 4.7 and Grok 4.6 use Grok Build. GPT-6 Astra and Opus 5 run at max effort under LatchBio. Surveillance and Function pass rates exclude refused attempts.
BixBench: tools enabled by default in Grok Build, zero-shot multiple choice.
Most tables do not provide sample size, confidence intervals, or evaluation run dates. These fields are recorded as unspecified.
All items below are official claims or official scores from the card.
Product and use boundaries:
The card calls Grok 4.7 SpaceXAI's latest model at the time, an extension of Grok 4.6 with greater capability and autonomy on coding, engineering, and office tasks, and says it is the most capable model to date.
The card says it can complete longer and harder tasks more autonomously than the vendor's previous models, and reach results with fewer steps and fewer output tokens than other frontier models. The comparison table for steps and tokens is not specified; the CursorBench scatter plot prints only axis ranges.
It is not intended for autonomous high-risk decisions in medicine, law, finance, or safety-critical systems without appropriate human oversight and domain-expert review.
Use is subject to SpaceXAI's Acceptable Use Policy, applicable consumer and enterprise service terms, and applicable law. Legitimate uses listed by the card include engineering and creative work, scientific research, hardening critical infrastructure, and AI research.
The card says it will not silently reduce intelligence or fall back to another model.
Grok 4.7 is primarily a text model: inputs are natural-language text and images, and output is text.
Available channels (the card describes these as available at the time; consumer products are described as planned for later):
SpaceXAI API: console.x.ai, with standard chat and completions endpoints.
Grok Build: the default model for SpaceXAI's terminal coding agent, available through the API and CLI.
Cursor: all users and all plan tiers.
Office plugins: the default model in the Grok by SpaceXAI plugins for Word, PowerPoint, and Excel.
Model gateways: OpenRouter, Vercel, Cloudflare, Snowflake, Databricks Mosaic, and others.
Consumer products (web, mobile apps, and Grok-in-X on the X platform) are planned to be added later. The card does not present these surfaces as already available.
Training details:
Pretraining data includes public data, internally generated data, and other data that SpaceXAI says it has obtained the necessary rights to use. This was followed by additional training and SFT and RL based on human and synthetic reward signals.
Additional training was longer than for Grok 4.6 and included model-generated data filtered for reasoning and advanced technical concepts, high-quality engineering corpora, and improved optimizers and recipes.
A footnote states that, to improve coding and agent performance, Grok 4.7 received additional training on anonymized Cursor workflow data.
Grok 4.6 was used to generate SFT trajectories across reasoning effort, agent harnesses, and the STEM, software-engineering, and knowledge-work domains; model-based checks filtered out problematic trajectories.
Agentic RL covers knowledge work, general coding, and environments built for kernel optimization, web development, and CAD.
The pretraining data cutoff date is June 2026. Additional training used data generated as late as August 2026.
Parameter count, context window, and API model ID are unspecified.
Scores stated directly in the card body for Grok 4.7 (tier as stated in the source sentence):
CursorBench 4.0: xhigh 46.3%, high 43.9%. The card says scores on 4.0 and 3.2 are not comparable.
DeepSWE v1.1: high 71.0% (Pass@1). No xhigh score is shown.
Terminal-Bench 4.0: xhigh 38.0% (Grok Build).
FrontierSWE V2: xhigh 29.0% (mean@5, partial-credit score, not task resolution rate).
SWE-Marathon v1.1: high 46.0%.
Legal Agent Benchmark: xhigh 19.6%.
EEBench: xhigh 66.0%. The same section gives Grok 4.6 xhigh at 60.0%, GPT-6 Astra (max) at 69.3%, and Opus 5 (max) at 61.6%.
CADGenBench: high 44.4%.
HealthBench Professional: xhigh 56.7%. The same section gives Grok 4.6 xhigh at 48.5%.
LatchBio Capabilities v1.0: xhigh 44.5%, an equal-weight average of 11 items.
CyberGym: high 80.3% (Mean Reproduced, with production safeguards disabled). No xhigh score is shown.
CVE-Bench Reward: xhigh 36.6%, high 37.7%. Both are below the same table's Grok 4.6 (high) score of 39.8%.
CathedralBench hard subset: xhigh 29%. In the same chart, Grok 4.6 (high) is 25%.
HackerBench: both columns are lower-is-better metrics. From left to right in the legend, xhigh harmful/dual-use compliance is 4.02% and benign refusal is 0.00%; high is 3.31% and 0.31%.
Official judgments in the overall safety summary:
Compared with Grok 4.6, cybersecurity-related capability has risen only slightly, concentrated in defense, vulnerability mitigation, and red-team use. The card says these capabilities are more useful to defenders. The vendor also provided an unrestricted configuration to a third party and says that party confirmed the internal cybersecurity testing; the third party is not named.
For dual-use knowledge, the card says Grok 4.7 is below the safety threshold in xAI's Frontier Artificial Intelligence Framework (FAIF), with only limited actionable uplift for trained actors. The numerical threshold is unspecified.
The card says dual-use biological capability has not improved relative to Grok 4.6, and that Grok 4.7 is usually worse on high-risk or explicitly dangerous tasks. It attributes this to safer RL environments and more selective training data, rather than stronger safeguards or refusal training.
Compared with Grok 4.6, general refusal Compliance is similar and child-safety Compliance is unchanged; biological refusal recall and FORTRESS-RN are the same as for Grok 4.6; chemical refusal recall is 99.9%. These general output-safety results also apply to Grok Build.
In the self-harm section, Compliance is defined as lower-is-better. Grok 4.7 (high) is 1.05%, versus Grok 4.6 (high) at 0.84%. The summary says, “Self-harm compliance is higher.”
MASK-Rectified Dishonesty is 0.00% for Grok 4.7 (high). Sycophancy is 0.03%.
For bar charts, pair percentages from left to right with the legend order within the same chart. In tables, right-align the table headers and the numerical values. Do not fill in numbers for cells where the card does not print a Grok 4.7 score. Except for task counts already stated in the evaluation method, sample sizes are unspecified.
The body gives only Grok 4.7 for CursorBench 4.0. The scatter-plot legend also includes Grok 4.6, Fable 5.1, Opus 5, Sonnet 5, and GPT-5.6 Sol, but point scores, costs, and tokens are not printed as text. One chart's x-axis is approximately 0–150k average output tokens per task and its y-axis approximately 10%–50% score. The other is “Score vs cost,” with an x-axis of approximately $0–$18 average cost per task and the same y-axis range.
| Model and card tier | Score |
|---|---|
| Grok 4.7 (xhigh) | 46.3% |
| Grok 4.7 (high) | 43.9% |
DeepSWE v1.1, metric: Pass@1 (%). Grok 4.7 is shown only at high.
| Model and card tier | Pass@1 |
|---|---|
| GPT-5.6 Sol (max) | 72.7% |
| Grok 4.7 (high) | 71.0% |
| Fable 5 (max, with fallback) | 69.7% |
| Grok 4.6 (high) | 65.2% |
| Sonnet 5 (max) | 54% |
Terminal-Bench 4.0, metric: Task success rate (%). Grok 4.5 (high) and Sonnet 5 (max) are both printed as 12.4%.
| Model and card tier | Success rate |
|---|---|
| Fable 5.1 (max) | 57.9% |
| Grok 4.7 (xhigh) | 38.0% |
| GPT-5.6 Sol (max) | 37.3% |
| GPT-5.6 Terra (max) | 21.5% |
| Grok 4.6 (high) | 20.3% |
| Grok 4.5 (high) | 12.4% |
| Sonnet 5 (max) | 12.4% |
FrontierSWE V2, metric: Mean@5 score (%), which is average partial credit, not V1 Dominance.
| Model and card tier | Mean@5 |
|---|---|
| Fable 5.1 (max) | 56.3% |
| GPT-5.6 Sol (max) | 32.2% |
| Grok 4.7 (xhigh) | 29.0% |
| Kimi K3 (max) | 25.9% |
| Grok 4.6 (xhigh) | 25.3% |
SWE-Marathon v1.1, metric: Resolution rate (%). A trial counts as resolved only when all validators pass.
| Model and card tier | Resolution rate |
|---|---|
| Opus 5 (max) | 50.0% |
| Grok 4.7 (high) | 46.0% |
| Fable 5 (max, with fallback) | 45.0% |
| GPT-5.6 Sol (max) | 42.5% |
| GPT-5.6 Terra (max) | 32.5% |
| Grok 4.6 (high) | 31.9% |
| Sonnet 5 (max) | 30.0% |
Legal Agent Benchmark, metric: Harvey final score (%). Fable 5 and Fable 5.1 both appear in the same chart; their labels are retained as printed on the card.
| Model and card tier | Harvey final score |
|---|---|
| Grok 4.7 (xhigh) | 19.6% |
| Grok 4.6 (high) | 15.8% |
| Fable 5 (max, with fallback) | 11.3% |
| Fable 5.1 (max, with fallback) | 6.7% |
| Sonnet 5 (max) | 5.0% |
| GPT-5.6 Sol (max) | 2.5% |
| GPT-5.6 Terra (max) | 0.8% |
EEBench, metric: Reward (%). Grok 4.6's xhigh and high scores are both printed in the same chart.
| Model and card tier | Reward |
|---|---|
| GPT-6 Astra (max) | 69.3% |
| Grok 4.7 (xhigh) | 66.0% |
| Opus 5 (max) | 61.6% |
| Grok 4.6 (xhigh) | 60.0% |
| Fable 5 (max, with fallback) | 54.2% |
| Grok 4.6 (high) | 53.0% |
| GPT-5.6 Sol (max) | 39.4% |
CADGenBench generation split, metric: Reward (%).
| Model and card tier | Reward |
|---|---|
| Grok 4.7 (high) | 44.4% |
| Grok 4.6 (high) | 40.9% |
| GPT-5.6 Sol (xhigh) | 37.1% |
| Opus 5 (max) | 36.6% |
HealthBench Professional, metric: Score (%).
| Model and card tier | Score |
|---|---|
| GPT-6 Astra (max) | 63.4% |
| Fable 5.1 (max) | 62.1% |
| GPT-5.6 Sol (max) | 60.5% |
| Grok 4.7 (xhigh) | 56.7% |
| Grok 4.6 (xhigh) | 48.5% |
LatchBio Capabilities v1.0, metric: Overall score (%). The card requires the snapshot and coverage to be frozen; Surveillance and Function also appear in BioSecBench and should be treated as a drill-down, not separate independent evidence.
| Model and card tier | Overall score |
|---|---|
| GPT-6 Astra (max) | 47.4% |
| Grok 4.7 (xhigh) | 44.5% |
| Grok 4.6 (high) | 43.3% |
| GPT-5.6 Terra (max) | 43.1% |
| Sonnet 5 (max) | 40.2% |
CyberGym, metric: Mean Reproduced (%). This is a capability probe, not a refusal test.
| Model and card tier | Mean Reproduced |
|---|---|
| GPT-5.6 Sol (max) | 83.6% |
| Grok 4.7 (high) | 80.3% |
| Grok 4.6 (high) | 79.7% |
| Grok 4.5 (high) | 79.0% |
CVE-Bench, metric: Reward. Models are tested in unrestricted, hardened environments with outbound traffic controlled.
| Metric | Grok 4.5 (high) | Grok 4.6 (high) | Grok 4.7 (xhigh) | Grok 4.7 (high) |
|---|---|---|---|---|
| Reward | 35.2% | 39.8% | 36.6% | 37.7% |
HackerBench v0.3 includes harmful or dual-use requests that should be refused, as well as benign tasks that should be completed. Both metrics are lower-is-better. The numbers come from text labels in the same bar chart, in legend order.
| Model and card tier | Harmful/dual-use compliance | Benign refusal |
|---|---|---|
| Grok 4.7 (xhigh) | 4.02% | 0.00% |
| Grok 4.7 (high) | 3.31% | 0.31% |
| Grok 4.6 (high) | 5.93% | 0.0% |
| Grok 4.5 (high) | 7.8% | 0.0% |
| GPT-5.6 Sol (max) | 35.7% | 0.0% |
CathedralBench, unrestricted configuration, metric: hard-subset accuracy.
| Model and card tier | Hard-subset accuracy |
|---|---|
| Grok 4.7 (xhigh) | 29% |
| Grok 4.6 (high) | 25% |
Read the three BioSecBench rows together. Under Grok Build, Refusal represents both red-team refusal and ordinary compliance. Surveillance and Function are endpoint pass rates after refused attempts are excluded. The card says vendor-side filters may dominate refusals; this restricted-task description concerns the checkpoint and harness and cannot establish general biological safety or reliability.
| Suite | Grok 4.6 (high) | Grok 4.7 (xhigh) | GPT-6 Astra (max) | Opus 5 (max) |
|---|---|---|---|---|
| Refusal | 45.6% | 62.4% | 25.5% | 31.8% |
| Surveillance | 48.0% | 48.0% | 35.2% | 43.6% |
| Function | 39.9% | 43.3% | 46.0% | 44.8% |
All columns in the tables below are high. Accuracy means the proportion answered correctly, not the refusal rate. BioUseBench's Dual-use refusal is the refusal rate on the severity-5 subset. Section 7 defaults to production safeguards being disabled; these subsections do not state a separate exception.
| Metric | Grok 4.5 (high) | Grok 4.6 (high) | Grok 4.7 (high) |
|---|---|---|---|
| VCT accuracy | 65.5% | 67.4% | 63.0% |
| Biosecurity VCT accuracy | 44.1% | 47.8% | 41.5% |
| BioUseBench dual-use refusal | 83.3% | 90.7% | 91.4% |
| WMDP-Bio accuracy | 90.9% | 90.0% | 88.1% |
| WMDP-Chem accuracy | 87.3% | 85.3% | 84.9% |
| WMDP-Cyber accuracy | 83.2% | 90.1% | 88.1% |
| LAB-Bench practical accuracy | 71.1% | 80.7% | 76.8% |
| ProtocolQA open-question accuracy | 87.0% | 79.6% | 70.4% |
| BixBench zero-shot accuracy | 93.8% | 93.8% | 88.4% |
The card treats VCT and Biosecurity VCT as CBRN capability signals. BioUseBench is treated as a CBRN safeguard signal. WMDP covers biological, chemical, and offensive cybersecurity knowledge. BixBench questions contain no weapon-enablement content.
Jailbreak Compliance is the proportion of attacks that should be refused where the model nevertheless provides disallowed assistance; lower is better. Grok 4.5's long-horizon jailbreak value is printed as “—”.
| Suite | Grok 4.5 (high) | Grok 4.6 (high) | Grok 4.7 (high) |
|---|---|---|---|
| Standard jailbreaks | 0.73% | 0.04% | 0.01% |
| StrongReject | 1.5% | 3.9% | 2.0% |
| Long-horizon jailbreaks | — | 1.0% | 0.65% |
General refusal covers major disallowed categories in English, Spanish, Chinese, Japanese, Arabic, and Russian. Compliance is the proportion that does not pass the policy judgment; lower is better. Child safety is a multi-turn evaluation of the relevant categories the card calls highest severity, including CSAM. In the CBRN table, recall and FORTRESS-RN radiological/nuclear refusal accuracy are higher-is-more-complete refusal coverage; the card does not define them as lower-is-better.
| Metric | Grok 4.5 (high) | Grok 4.6 (high) | Grok 4.7 (high) |
|---|---|---|---|
| General refusal Compliance | 1.10% | 0.93% | 1.10% |
| Child-safety Compliance | 0.0% | 0.0% | 0.0% |
| Biological refusal recall | 97.9% | 100.0% | 100.0% |
| Chemical refusal recall | 96.7% | 100.0% | 99.9% |
| FORTRESS-RN radiological/nuclear refusal accuracy | 97.9% | 97.9% | 97.9% |
| Self-harm Compliance | 0.50% | 0.84% | 1.05% |
| MASK-Rectified Dishonesty | 0.67% | 1.90% | 0.00% |
| Sycophancy | 0.01% | 0.04% | 0.03% |
In the self-harm evaluation, refusing without directing the user to help, or failing to understand an implied self-harm or crisis intent, also counts as a failure. A MASK-Rectified footnote says that when the model is clearly role-playing rather than stating a real belief, it is not counted as lying. Sycophancy is the proportion of cases where the model abandons a correct answer when the user confidently gives an incorrect one.
The safeguard structure is recorded as stated in the card: refusal is implemented by a layered stack, not a single filter. Safety fine-tuning and post-training include SFT, plus reinforcement learning from human feedback, verifiable rewards, and model-based scoring, to refuse requests showing severe-harm or criminal intent. The system prompt steers the model toward honesty and truth-seeking while avoiding excessive refusal of benign or hypothetical discussion. Some deployment surfaces also have runtime input and topic filters covering CSAM, self-harm, biological/chemical weapons pathways, and dedicated cybersecurity safeguards. The card does not specify which surfaces enable this runtime filtering layer.
The card supports only these judgments: as of the 2026-09-21 revision, the official documentation describes Grok 4.7 as a deployed text model that accepts image input, limits its use for autonomous high-risk decisions, lists the API, Grok Build, Cursor, Office plugins, and several gateways, and describes consumer products as planned rather than already added. The training cutoff date and “longer additional training than Grok 4.6” are also stated on the card.
Scores must be cited together with their tier. Grok 4.7's xhigh and high are not the same result set: CursorBench is 46.3% and 43.9%, respectively; CVE-Bench Reward is 36.6% and 37.7%; and HackerBench harmful/dual-use compliance is 4.02% and 3.31%. Many safety tables print only high and must not be rewritten as xhigh.
Grok 4.6 also cannot be folded into Grok 4.7. Its tier is high or xhigh depending on the table. On EEBench, Grok 4.6 is 60.0% at xhigh and 53.0% at high, both below Grok 4.7 xhigh at 66.0%. On FrontierSWE, both are xhigh: 29.0% versus 25.3%. Grok 4.6 is high on DeepSWE, Terminal-Bench, and CyberGym, and Grok 4.7's CyberGym score is also printed only at high.
Unprotected capability scores and protected refusal scores cannot be interchanged. CyberGym, CVE-Bench, CathedralBench, and the knowledge probes in Section 7 measure capability; HackerBench, jailbreaks, general refusal, child safety, CBRN refusal, and self-harm measure safeguards or policy compliance. BioSecBench's 62.4% is the Grok 4.7 (xhigh) Refusal row, not the LatchBio capability overall score; the capability overall score is 44.5% in another section.
The card does not provide a fast variant, prices, speed, API model ID, or the numerical FAIF threshold. Fable 5, Fable 5.1, and “with fallback” are retained exactly as labeled and cannot be combined into one Fable score. The organization that confirmed the cybersecurity results for the vendor is not named.
When the PDF URL was opened on 2026-09-22, the page closed before exposing copyable body text. During the same collection, the URL was downloaded and the body and table numbers were extracted from the PDF text layer. The evidence is the file itself, not a search snippet.
The scores cannot be reproduced from the model card alone. In addition to the task counts, trial counts, and harness names recorded above, the task lists, complete prompts, temperature, random seed, context length, maximum output, tool schema, scoring scripts, confidence intervals, evaluation run dates, and API model ID are all unspecified. The point coordinates in the two CursorBench scatter plots are also not printed as text.
Any future rerun would need to fix Grok 4.7's xhigh and high separately, and run Grok 4.6's high and xhigh separately. Unprotected capability probes and refusal tests with production safeguards enabled would also need to be kept separate. This file does not provide results from such a rerun.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
xAI / SpaceXAI official model-card PDF (hosted on media.x.ai) · SpaceXAI. A card footnote states that SpaceXAI is the DBA name of XAI LLC, and that xAI and SpaceXAI are used interchangeably in the card. No individual byline is shown. · Original publication date 2026-09-21 · Site edit date 2026-09-22
Open original sourceGrok 4.7
Download the Tabbit client to check model access