Grok 4.7 is the model to pilot when a task is long, tool-heavy, and worth a larger token bill. Grok 4.6 remains the default when the work is routine and you want the shorter trace. The API sticker did not change.
xAI published Grok 4.7 on September 21, 2026, and said it is served at the same price and speed as Grok 4.6. The model cards opened on September 22 agree on the rate card and the 500K context window. They do not agree that a finished task costs the same. Artificial Analysis, evaluating Grok 4.7 at xhigh, counted about 81,000 output tokens per Intelligence Index task, against about 36,000 for Grok 4.6 at high. That is the number that decides the switch. (xAI launch, Grok 4.7 model card, Grok 4.6 model card, Artificial Analysis)
The same tension showed up immediately in Cursor. On September 21, u/Dynamix86 wrote that one Ultra account, extra high, Fast off, burned plan percentage about 2.5 times faster on Grok 4.7 than on Grok 4.6 for the same unnamed task, and that most work would go back to 4.6.

That post is one account and one task. It does not prove API invoices or answer quality. It does show why "same price" is the wrong first question. The Grok 4.6 overview covers the previous model on its own. This page is only the choice between the two. Specs for each model also live on the Grok 4.7 model page and the Grok 4.6 model page.
Key takeaways
Both API cards list $2 / $0.50 cached / $6 per million tokens under 200K context, and $4 / $1 / $12 above 200K. The 500K window matches.
Grok 4.7 xhigh used about 81k output tokens per Artificial Analysis intelligence task, versus about 36k for Grok 4.6 high and about 38k for Grok 4.6 xhigh. Output-only math at $6 is about $0.49 versus $0.22–$0.23.
The Intelligence Index move is +2, and it compares xhigh with high (46 versus 44). The Coding Agent Index move is +9 at matched xhigh inside Grok Build (56 versus 47).
xAI's launch table improves on every listed row, but the columns are Grok 4.7 xHigh and Grok 4.6 High. DeepSWE on that table is marked high effort for 4.7. Those gaps are not a same-rung experiment.
Keep Grok 4.6 for short, repeated work. Pilot Grok 4.7 when a long coding run or a document-heavy task fails on 4.6 often enough to pay for the extra tokens.
No same-task Tabbit run was completed for this article. A model page is not a measured win.
The one number: same $2/$6 rate, about 81k versus 36k output tokens
The launch headline says Grok 4.7 is "twice as fast, at half the price of comparable models," and the body says it is served at the same price and speed as Grok 4.6. The first sentence does not name the comparable models or the speed unit. The second sentence is the one the rate cards support.
What changed is how many tokens the newer model spends. Artificial Analysis wrote that Grok 4.7 xhigh uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 high, which they call 125% more. Later in the same article they say 81k is more than double the 38k used by Grok 4.6 xhigh. Both comparisons matter. The 36k figure is not a same-effort pair. The 38k figure is.
Output tokens per Intelligence Index task (Artificial Analysis, Sept 21, 2026)
Grok 4.7 xhigh ~81k
Grok 4.6 high ~36k (+125% tokens at a higher effort label)
Grok 4.6 xhigh ~38k (more than 2× tokens at the same effort label)
Output-only charge at the shared $6 / 1M rate
81k × $6 / 1M ≈ $0.49
36k × $6 / 1M ≈ $0.22
38k × $6 / 1M ≈ $0.23That arithmetic is not an invoice. It ignores input tokens, cache hits, tool calls, retries, and any client that bills a plan percentage instead of the API. It also ignores the score that those tokens bought.
On the Intelligence Index the score moved from 44 for Grok 4.6 high to 46 for Grok 4.7 xhigh. Outside agentic knowledge work, Artificial Analysis said Grok 4.7 broadly matched Grok 4.6 high, with Terminal-Bench 4.0 up 4.5 points and GDP.pdf up 3.0 points, and with regressions on AA-LCR (−3.7 points) and AutomationBench-AA (−1.1 points). Absolute scores for those four rows were not printed in the article text.
The larger movement is the coding-agent index, and it is a same-effort comparison. Grok 4.7 xhigh with Grok Build scored 56, up from 47 for Grok 4.6 xhigh with Grok Build. Inside that harness, DeepSWE v1.1 went from 65% to 73%, Terminal-Bench 4.0 from 18% to 33%, and SWE-Atlas-QnA from 58% to 63%. Those terminal numbers are not the 38.0% and 20.3% on xAI's launch table, and they are not the +4.5 point unified-harness delta above. Three different Terminal-Bench readings can all be real if the harness and effort differ.
Use Grok 4.7 when the extra tokens are buying a long run you would otherwise redo. Stay on Grok 4.6 when a +2 index point, measured across effort labels, is not worth roughly twice the output.
Specifications and price at a glance
Checked September 22, 2026, on the two model cards. "Higher context" on those cards means a request above 200K tokens. Cursor's own docs use a 256K boundary and are not this table.
| Dimension | Grok 4.7 | Grok 4.6 | How to use the row |
|---|---|---|---|
| API model ID | grok-4.7 | grok-4.6 | Pin the ID. Do not rely on a "latest" alias for a comparison. |
| Context window | 500,000 | 500,000 | A client can expose less. Cursor's Grok 4.7 page describes 256K standard and 500K long context. |
| Modalities | Text and image in, text out | Text and image in, text out | Image size limits are on the shared models docs, not a reason to pick one ID. |
| Input / cached / output, at or under 200K | $2 / $0.50 / $6 per 1M | $2 / $0.50 / $6 per 1M | Same list rate. |
| Input / cached / output, above 200K | $4 / $1 / $12 per 1M | $4 / $1 / $12 per 1M | The long-context multiplier matches too. |
| Knowledge cutoff | May 2026, on the models index | Not restated on the model card opened this date | Do not copy an older cutoff forward. |
| Fast option | Launch post: a fast variant at twice the output speed and twice the price, with no benchmark row | Not stated on the card opened this date | Do not assume the two Fast switches are the same product. |
xAI's launch table also prints $2 input and $6 output for both Grok columns, next to GPT-5.6 Sol at $4 / $20 and Fable 5.1 at $10 / $50. That table is a list-price comparison. It is not a completed-task comparison. Fable 5.1 still leads several of xAI's own rows; the Fable 5.1 review is the place for that model's trade-offs, not a claim that Grok replaced it.
Cursor's Grok 4.7 documentation, opened the same day and served in Chinese at cursor.com/cn/docs/models/grok-4-7, puts Grok 4.7 in a usage pool with Grok 4.6, Grok 4.5, and Composer 2.5. On-demand rates shown there match the short-context API card: $2 input, $0.50 cached, $6 output. Fast is listed at $4 / $1 / $12, and the page says Fast is the default speed on Pro and above. Input past 256K is billed at 2× for standard and 3× the standard rate for Fast, up to 500K. A Cursor percentage and an API dollar are different meters. Compare them only after you fix effort and Fast on both sides.
Benchmarks, and how much to trust them
Two tables below. The first is xAI's launch table, with the effort labels the page actually printed. The second is Artificial Analysis where the effort labels match, plus the index scores that do not.
| xAI launch table, September 21, 2026 | Grok 4.7 | Grok 4.6 | GPT-5.6 Sol | Fable 5.1 |
|---|---|---|---|---|
| Effort label on the column | xHigh; DeepSWE marked high effort | High | Max | Max |
| List price, input / output per 1M | $2 / $6 | $2 / $6 | $4 / $20 | $10 / $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0% | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
Grok 4.7 is ahead of Grok 4.6 on every row of that table. It is not ahead of Fable 5.1 on CursorBench, AA Briefcase, Terminal-Bench, or HealthBench Professional. It is ahead of GPT-5.6 Sol on five of the seven score rows and behind on DeepSWE and HealthBench. Sample size, harness, and confidence intervals are not on the page. The chart caption says benchmarks are directional vendor results.
The CursorBench scatter on the same page does not include Grok 4.6. For Grok 4.7 it shows a steep effort ladder: Extra High 46.3% at about $6.01 and 70,141 output tokens per task; High 43.9% at $4.69; Medium 41.6% at $3.49; Low 33.1% at $1.58. "Grok 4.7" without an effort label is not one score.
| Artificial Analysis, article dated September 21, 2026 | Grok 4.7 | Grok 4.6 | Effort pair | What it is safe to say |
|---|---|---|---|---|
| Intelligence Index v4.3 | 46 | 44 | xhigh vs high | +2 points, not a same-rung result |
| Output tokens per intelligence task | ~81k | ~36k high; ~38k xhigh | mixed | The bill moves more than the index |
| Coding Agent Index, Grok Build | 56 | 47 | xhigh vs xhigh | +9 on the vendor coding agent |
| DeepSWE v1.1, Grok Build | 73% | 65% | xhigh vs xhigh | Not the 71.0 / 65.2 launch-table pair |
| Terminal-Bench 4.0, Grok Build | 33% | 18% | xhigh vs xhigh | Not xAI's 38.0 / 20.3 pair |
| SWE-Atlas-QnA, Grok Build | 63% | 58% | xhigh vs xhigh | Narrower than the terminal jump |
| AA-Briefcase Elo | 1,657 | 1,546 at high | xhigh vs high | Analytical quality up, presentation down |
| GDPval-AA Elo | 1,695 | 1,605 at high | xhigh vs high | Knowledge-work gain, still behind Fable 5.1's 1,735 on xAI's chart |
| Hallucination rate, AA-Omniscience | 29% | 34% at high | xhigh vs high | Accuracy stayed 47% vs 48% |
Read the coding-agent block when you care about Grok Build. Read the intelligence block when you care about a shared harness. Do not average them into one "Grok is 10% better" line. A Hacker News comment asked the obvious question: why compare 4.7 xhigh with 4.6 high? Replies on that thread disagree about when xhigh arrived for 4.6. Artificial Analysis does publish a Grok 4.6 xhigh coding-agent score, so the label exists in at least one current eval. The launch table still does not use it.

Where Grok 4.7 actually pulls ahead
Long coding runs inside a coding agent
The cleanest gain is the one with matched effort. On Grok Build, the coding-agent index rose 9 points, and every component moved up. Terminal work moved the most, from 18% to 33% in that harness. If your failure mode is a multi-hour repository task that stalls, this is the row that justifies a pilot. It does not say Grok Build will beat Claude Code or Codex. Artificial Analysis ranked Grok 4.7 plus Grok Build fourth among native harnesses, behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.
xAI's own Terminal-Bench 4.0 jump, 20.3% to 38.0%, points the same direction and uses unmatched effort labels. Treat it as supporting vendor evidence, not as a second independent measurement of the 18-to-33 result.
Documents and analysis, with a presentation caveat
On AA-Briefcase, Grok 4.7 scored 1,657 Elo, 111 above Grok 4.6 high, just behind Claude Opus 5 and Claude Fable 5.1 in Artificial Analysis's write-up. Analytical quality was 1,994 Elo versus 1,690. Presentation quality was 1,499 versus 1,519. The model got better at the analysis and slightly worse at the deck.
Artificial Analysis said the same split more sharply on the public AA-Briefcase-Lite due-diligence scenario: analytical Elo 1,698 to 1,994, presentation Elo 1,531 to 1,499. Example decks cost about $8 for Grok 4.7 xhigh and about $4.40 for Grok 4.6 xhigh. That pair is same-effort and about 1.8× the API cost for a stronger analysis and a weaker presentation. It is one scenario, not a universal office benchmark. GDPval-AA moved from 1,605 to 1,695 Elo on the high-versus-xhigh pair, and xAI's GDPval chart puts Fable 5.1 max at 1,735.

A lower hallucination rate, not a higher accuracy rate
AA-Omniscience hallucination rate was 29% for Grok 4.7 xhigh and 34% for Grok 4.6 high. Accuracy was 47% versus 48%. The index moved from 30 to 32. If you need fewer confident wrong answers, that is a real shift. If you need more correct answers, this pair does not show it. The effort labels still differ.
Where Grok 4.6 is still the better default
Short work, where the extra thinking is the product
Paying the same rate for about twice the output is a good trade only when the extra output prevents a retry. For a bounded extraction, a short patch, or a status JSON, Grok 4.6's shorter trace is the feature. Hacker News readers looking at the intelligence-token chart called the change a regression in token efficiency. That comment does not add a new measurement. It is the same Artificial Analysis chart, read as a cost problem rather than a capability upgrade.

On the tasks Artificial Analysis called out as regressions versus Grok 4.6 high, AA-LCR and AutomationBench-AA, there is no reason to leave 4.6. Absolute scores were not published, so the size of those drops is the published delta only: 3.7 and 1.1 points.
Cursor plans, if Fast and effort are not pinned
A list rate of $2/$6 does not describe a Cursor Ultra percentage. Dynamix86's timings are a single-user log: Grok 4.7 intervals of 34, 55, and 38 minutes per percentage point, against 133 and 101 minutes on Grok 4.6 earlier the same day, both at extra high with Fast off. The author also read an Artificial Analysis chart as $8.82 versus $3.53 per coding-agent task. Those dollar figures were the author's reading of a chart. They were not printed as a sentence in the Artificial Analysis article opened for this page, so they stay attributed to the post.
A reply on the same thread was blunter: the release looked poor because the cost was clearly too high even if benchmarks were set aside.

Cursor's docs also say Fast is the default on Pro and higher, at double the token rates. If one side of a personal test is Fast and the other is not, the 2× price is a setting, not a model change. The Reddit log says Fast was off. Many casual comparisons will not.
Anything you already accept from Grok 4.6
Grok 4.6's review notes already describe a 500K agent model at this rate. Grok 4.7 does not add a new context tier on the API card, and it does not cut the list price. If 4.6 already finishes the job and you are not blocked on long terminal work or long documents, the upgrade is optional. HealthBench Professional still trails GPT-5.6 Sol and Fable 5.1 on xAI's own table (56.7 versus 60.5 and 62.1), so a clinical-reasoning switch is not the story either.
What people actually said
The early comments split into a cost camp and a feel camp. Both are hours old. Neither camp published a prompt, a repository, or a score.
The bill. Dynamix86's 2.5× plan-usage estimate and truecakesnake's "cost is clearly too high" sit next to the token chart. A separate Hacker News commenter, notduckrabbit, pointed at the same efficiency gap. Together they say: do not migrate a high-volume route on launch day.
The feel. Other r/cursor comments, posted the same afternoon, liked the model and did not mention price.



The speed and "does not overthink" comments can both be true for a chat session and still false for an xhigh coding agent that emits 81k tokens. "Not as slow as 4.6" was not paired with a token count. Artificial Analysis measured about 188 tokens per second on long prompts and about 7.1 minutes of decode time per intelligence task, excluding time to first token and overhead. A fast first impression and a long agent trace are compatible.
YouTube search for "Grok 4.7 vs Grok 4.6" on September 22 returned launch explainers, including pre-release videos that guess parameter counts. Those counts are not on the xAI launch page, which only says "a new, larger base model." No video comment is used here.
The verdict
There is no overall winner. Pick by the shape of the work, and keep the effort label in the decision.
| Workload | Choose | Why | What to watch |
|---|---|---|---|
| Short extraction, classification, or a small patch | Grok 4.6 | The index gain is small and the output trace is much shorter | Do not pay xhigh for a one-line answer |
| Repeated Cursor work on a usage pool | Grok 4.6, unless a task is failing | One Ultra log showed about 2.5× plan usage at extra high with Fast off | Pin Fast and effort before blaming the model |
| Multi-hour coding in Grok Build or a similar agent | Pilot Grok 4.7 xhigh | Coding Agent Index 56 vs 47 at matched xhigh | Compare completed tasks, not the index alone |
| Market models, memos, and target decks | Pilot Grok 4.7, then edit the deck | Analytical Elo jumped; presentation Elo slipped; example decks were about $8 vs $4.40 | Budget a human pass on slides |
| Long-context API calls above 200K | Either, on price | Both cards double the rates past 200K | Cursor's 256K boundary is a different meter |
| You need Fable-level CursorBench or Terminal-Bench | Not this pair | Fable 5.1 Max is 51.8% and 57.9% on xAI's table | Price that choice on Fable's card, not Grok's |
A practical test is one accepted task, same effort, Fast off, same harness. Record completed or not, human edits, output tokens, and either API dollars or plan percentage. One success is not a rate.
Run the same task in Tabbit before you migrate a route
A benchmark row does not tell you how either model behaves on your tabs, docs, and half-finished research. Tabbit Browser is the workspace for that check: the model can sit next to the pages you are already reading, instead of in a separate chat that never sees them. The agentic browser guide and the agentic reasoning guide separate a model score from a finished multi-step workflow. The AI browser comparison and the Tabbit browser guide cover the product around the model. The 2026 AI browser roundup is the wider set if Tabbit is not the client you want.

Both Grok IDs are registered for Tabbit. Whether Grok 4.7 or Grok 4.6 appears in your picker depends on the account and the current selector. This article does not include a Tabbit transcript, a token log, or a winner. If both IDs are available, run one task you already know how to grade, keep effort and tools the same, and keep Grok 4.6 if the only change is a longer trace. The Tabbit practices guide is about the browser workflow, not a claim that either Grok build is preinstalled for every account.
API list price, a Cursor pool, and Tabbit access are three different bills. Do not add them together.
FAQ
Is Grok 4.7 better than Grok 4.6?
It depends on the task and the reasoning effort. On Artificial Analysis's Coding Agent Index, Grok 4.7 xhigh with Grok Build scored 56 versus 47 for Grok 4.6 xhigh. The Intelligence Index only moved from 44 at Grok 4.6 high to 46 at Grok 4.7 xhigh, and some non-agent tasks regressed. There is no single winner.
Do Grok 4.7 and Grok 4.6 cost the same?
On the xAI model cards opened September 22, 2026, both list $2 per million input tokens, $0.50 cached input, and $6 output below 200K context, then $4, $1, and $12 above 200K. Cursor documents a separate usage pool and a Fast tier. The list rate matching does not mean a finished task costs the same.
Why can Grok 4.7 cost more if the token price is unchanged?
Artificial Analysis measured about 81,000 output tokens per Intelligence Index task for Grok 4.7 xhigh, versus about 36,000 for Grok 4.6 high and about 38,000 for Grok 4.6 xhigh. At $6 per million output tokens, that output alone is roughly $0.49 versus $0.22 or $0.23, before input, cache, tools, or retries.
Which Grok 4.7 benchmark gains are on the same effort level?
The Coding Agent Index comparison is xhigh versus xhigh: 56 versus 47, with DeepSWE 73% versus 65%, Terminal-Bench 4.0 33% versus 18%, and SWE-Atlas-QnA 63% versus 58% inside Grok Build. xAI's launch table compares Grok 4.7 xHigh with Grok 4.6 High, and DeepSWE on that table is marked high effort for Grok 4.7. Do not treat those rows as a same-rung test.
Should I switch from Grok 4.6 to Grok 4.7 in Cursor?
Switch a long coding or document task if you can accept a larger usage drop. One Cursor Ultra user, on extra high and with Fast off, estimated about 2.5 times the plan usage and planned to move most work back to 4.6. Other early comments said 4.7 felt faster or less prone to overthinking. Pin the effort and Fast setting before you compare.
Can I compare Grok 4.7 and Grok 4.6 in Tabbit Browser?
Both models have Tabbit model pages. Whether they appear in your selector depends on the account and the live picker. This article does not include a same-task Tabbit run, so a selector entry is not evidence that one model wins a workflow.