TabbitBlog

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

In this article
  1. Key takeaways
  2. The one number: same $2/$6 rate, about 81k versus 36k output tokens
  3. Specifications and price at a glance
  4. Benchmarks, and how much to trust them
  5. Where Grok 4.7 actually pulls ahead
  6. Long coding runs inside a coding agent
  7. Documents and analysis, with a presentation caveat
  8. A lower hallucination rate, not a higher accuracy rate
  9. Where Grok 4.6 is still the better default
  10. Short work, where the extra thinking is the product
  11. Cursor plans, if Fast and effort are not pinned
  12. Anything you already accept from Grok 4.6
  13. What people actually said
  14. The verdict
  15. Run the same task in Tabbit before you migrate a route

Grok 4.7 is the model to pilot when a task is long, tool-heavy, and worth a larger token bill. Grok 4.6 remains the default when the work is routine and you want the shorter trace. The API sticker did not change.

xAI published Grok 4.7 on September 21, 2026, and said it is served at the same price and speed as Grok 4.6. The model cards opened on September 22 agree on the rate card and the 500K context window. They do not agree that a finished task costs the same. Artificial Analysis, evaluating Grok 4.7 at xhigh, counted about 81,000 output tokens per Intelligence Index task, against about 36,000 for Grok 4.6 at high. That is the number that decides the switch. (xAI launch, Grok 4.7 model card, Grok 4.6 model card, Artificial Analysis)

The same tension showed up immediately in Cursor. On September 21, u/Dynamix86 wrote that one Ultra account, extra high, Fast off, burned plan percentage about 2.5 times faster on Grok 4.7 than on Grok 4.6 for the same unnamed task, and that most work would go back to 4.6.

Reddit post by Dynamix86 comparing Cursor Ultra usage of Grok 4.7 and Grok 4.6
u/Dynamix86 on r/cursor, September 21, 2026: same task, extra high, Fast off, and a plan-usage estimate of about 2.5 times.

That post is one account and one task. It does not prove API invoices or answer quality. It does show why "same price" is the wrong first question. The Grok 4.6 overview covers the previous model on its own. This page is only the choice between the two. Specs for each model also live on the Grok 4.7 model page and the Grok 4.6 model page.

Key takeaways

  • Both API cards list $2 / $0.50 cached / $6 per million tokens under 200K context, and $4 / $1 / $12 above 200K. The 500K window matches.

  • Grok 4.7 xhigh used about 81k output tokens per Artificial Analysis intelligence task, versus about 36k for Grok 4.6 high and about 38k for Grok 4.6 xhigh. Output-only math at $6 is about $0.49 versus $0.22–$0.23.

  • The Intelligence Index move is +2, and it compares xhigh with high (46 versus 44). The Coding Agent Index move is +9 at matched xhigh inside Grok Build (56 versus 47).

  • xAI's launch table improves on every listed row, but the columns are Grok 4.7 xHigh and Grok 4.6 High. DeepSWE on that table is marked high effort for 4.7. Those gaps are not a same-rung experiment.

  • Keep Grok 4.6 for short, repeated work. Pilot Grok 4.7 when a long coding run or a document-heavy task fails on 4.6 often enough to pay for the extra tokens.

  • No same-task Tabbit run was completed for this article. A model page is not a measured win.

The one number: same $2/$6 rate, about 81k versus 36k output tokens

The launch headline says Grok 4.7 is "twice as fast, at half the price of comparable models," and the body says it is served at the same price and speed as Grok 4.6. The first sentence does not name the comparable models or the speed unit. The second sentence is the one the rate cards support.

What changed is how many tokens the newer model spends. Artificial Analysis wrote that Grok 4.7 xhigh uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 high, which they call 125% more. Later in the same article they say 81k is more than double the 38k used by Grok 4.6 xhigh. Both comparisons matter. The 36k figure is not a same-effort pair. The 38k figure is.

Output tokens per Intelligence Index task (Artificial Analysis, Sept 21, 2026)
Grok 4.7 xhigh     ~81k
Grok 4.6 high      ~36k     (+125% tokens at a higher effort label)
Grok 4.6 xhigh     ~38k     (more than 2× tokens at the same effort label)

Output-only charge at the shared $6 / 1M rate
81k × $6 / 1M  ≈  $0.49
36k × $6 / 1M  ≈  $0.22
38k × $6 / 1M  ≈  $0.23

That arithmetic is not an invoice. It ignores input tokens, cache hits, tool calls, retries, and any client that bills a plan percentage instead of the API. It also ignores the score that those tokens bought.

On the Intelligence Index the score moved from 44 for Grok 4.6 high to 46 for Grok 4.7 xhigh. Outside agentic knowledge work, Artificial Analysis said Grok 4.7 broadly matched Grok 4.6 high, with Terminal-Bench 4.0 up 4.5 points and GDP.pdf up 3.0 points, and with regressions on AA-LCR (−3.7 points) and AutomationBench-AA (−1.1 points). Absolute scores for those four rows were not printed in the article text.

The larger movement is the coding-agent index, and it is a same-effort comparison. Grok 4.7 xhigh with Grok Build scored 56, up from 47 for Grok 4.6 xhigh with Grok Build. Inside that harness, DeepSWE v1.1 went from 65% to 73%, Terminal-Bench 4.0 from 18% to 33%, and SWE-Atlas-QnA from 58% to 63%. Those terminal numbers are not the 38.0% and 20.3% on xAI's launch table, and they are not the +4.5 point unified-harness delta above. Three different Terminal-Bench readings can all be real if the harness and effort differ.

Use Grok 4.7 when the extra tokens are buying a long run you would otherwise redo. Stay on Grok 4.6 when a +2 index point, measured across effort labels, is not worth roughly twice the output.

Specifications and price at a glance

Checked September 22, 2026, on the two model cards. "Higher context" on those cards means a request above 200K tokens. Cursor's own docs use a 256K boundary and are not this table.

DimensionGrok 4.7Grok 4.6How to use the row
API model IDgrok-4.7grok-4.6Pin the ID. Do not rely on a "latest" alias for a comparison.
Context window500,000500,000A client can expose less. Cursor's Grok 4.7 page describes 256K standard and 500K long context.
ModalitiesText and image in, text outText and image in, text outImage size limits are on the shared models docs, not a reason to pick one ID.
Input / cached / output, at or under 200K$2 / $0.50 / $6 per 1M$2 / $0.50 / $6 per 1MSame list rate.
Input / cached / output, above 200K$4 / $1 / $12 per 1M$4 / $1 / $12 per 1MThe long-context multiplier matches too.
Knowledge cutoffMay 2026, on the models indexNot restated on the model card opened this dateDo not copy an older cutoff forward.
Fast optionLaunch post: a fast variant at twice the output speed and twice the price, with no benchmark rowNot stated on the card opened this dateDo not assume the two Fast switches are the same product.

xAI's launch table also prints $2 input and $6 output for both Grok columns, next to GPT-5.6 Sol at $4 / $20 and Fable 5.1 at $10 / $50. That table is a list-price comparison. It is not a completed-task comparison. Fable 5.1 still leads several of xAI's own rows; the Fable 5.1 review is the place for that model's trade-offs, not a claim that Grok replaced it.

Cursor's Grok 4.7 documentation, opened the same day and served in Chinese at cursor.com/cn/docs/models/grok-4-7, puts Grok 4.7 in a usage pool with Grok 4.6, Grok 4.5, and Composer 2.5. On-demand rates shown there match the short-context API card: $2 input, $0.50 cached, $6 output. Fast is listed at $4 / $1 / $12, and the page says Fast is the default speed on Pro and above. Input past 256K is billed at 2× for standard and 3× the standard rate for Fast, up to 500K. A Cursor percentage and an API dollar are different meters. Compare them only after you fix effort and Fast on both sides.

Benchmarks, and how much to trust them

Two tables below. The first is xAI's launch table, with the effort labels the page actually printed. The second is Artificial Analysis where the effort labels match, plus the index scores that do not.

xAI launch table, September 21, 2026Grok 4.7Grok 4.6GPT-5.6 SolFable 5.1
Effort label on the columnxHigh; DeepSWE marked high effortHighMaxMax
List price, input / output per 1M$2 / $6$2 / $6$4 / $20$10 / $50
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0%65.2%72.7%70.0%
EEBench64.0%53.0%39.4%56.4%
AA Briefcase v1.11,6571,5461,4871,678
Terminal-Bench 4.038.0%20.3%37.3%57.9%
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
HealthBench Professional56.7%48.5%60.5%62.1%

Grok 4.7 is ahead of Grok 4.6 on every row of that table. It is not ahead of Fable 5.1 on CursorBench, AA Briefcase, Terminal-Bench, or HealthBench Professional. It is ahead of GPT-5.6 Sol on five of the seven score rows and behind on DeepSWE and HealthBench. Sample size, harness, and confidence intervals are not on the page. The chart caption says benchmarks are directional vendor results.

The CursorBench scatter on the same page does not include Grok 4.6. For Grok 4.7 it shows a steep effort ladder: Extra High 46.3% at about $6.01 and 70,141 output tokens per task; High 43.9% at $4.69; Medium 41.6% at $3.49; Low 33.1% at $1.58. "Grok 4.7" without an effort label is not one score.

Artificial Analysis, article dated September 21, 2026Grok 4.7Grok 4.6Effort pairWhat it is safe to say
Intelligence Index v4.34644xhigh vs high+2 points, not a same-rung result
Output tokens per intelligence task~81k~36k high; ~38k xhighmixedThe bill moves more than the index
Coding Agent Index, Grok Build5647xhigh vs xhigh+9 on the vendor coding agent
DeepSWE v1.1, Grok Build73%65%xhigh vs xhighNot the 71.0 / 65.2 launch-table pair
Terminal-Bench 4.0, Grok Build33%18%xhigh vs xhighNot xAI's 38.0 / 20.3 pair
SWE-Atlas-QnA, Grok Build63%58%xhigh vs xhighNarrower than the terminal jump
AA-Briefcase Elo1,6571,546 at highxhigh vs highAnalytical quality up, presentation down
GDPval-AA Elo1,6951,605 at highxhigh vs highKnowledge-work gain, still behind Fable 5.1's 1,735 on xAI's chart
Hallucination rate, AA-Omniscience29%34% at highxhigh vs highAccuracy stayed 47% vs 48%

Read the coding-agent block when you care about Grok Build. Read the intelligence block when you care about a shared harness. Do not average them into one "Grok is 10% better" line. A Hacker News comment asked the obvious question: why compare 4.7 xhigh with 4.6 high? Replies on that thread disagree about when xhigh arrived for 4.6. Artificial Analysis does publish a Grok 4.6 xhigh coding-agent score, so the label exists in at least one current eval. The launch table still does not use it.

Hacker News comment asking why Grok 4.7 xhigh is compared with Grok 4.6 high
AM1010101 on Hacker News, September 22, 2026: the launch comparison mixes effort labels.

Where Grok 4.7 actually pulls ahead

Long coding runs inside a coding agent

The cleanest gain is the one with matched effort. On Grok Build, the coding-agent index rose 9 points, and every component moved up. Terminal work moved the most, from 18% to 33% in that harness. If your failure mode is a multi-hour repository task that stalls, this is the row that justifies a pilot. It does not say Grok Build will beat Claude Code or Codex. Artificial Analysis ranked Grok 4.7 plus Grok Build fourth among native harnesses, behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.

xAI's own Terminal-Bench 4.0 jump, 20.3% to 38.0%, points the same direction and uses unmatched effort labels. Treat it as supporting vendor evidence, not as a second independent measurement of the 18-to-33 result.

Documents and analysis, with a presentation caveat

On AA-Briefcase, Grok 4.7 scored 1,657 Elo, 111 above Grok 4.6 high, just behind Claude Opus 5 and Claude Fable 5.1 in Artificial Analysis's write-up. Analytical quality was 1,994 Elo versus 1,690. Presentation quality was 1,499 versus 1,519. The model got better at the analysis and slightly worse at the deck.

Artificial Analysis said the same split more sharply on the public AA-Briefcase-Lite due-diligence scenario: analytical Elo 1,698 to 1,994, presentation Elo 1,531 to 1,499. Example decks cost about $8 for Grok 4.7 xhigh and about $4.40 for Grok 4.6 xhigh. That pair is same-effort and about 1.8× the API cost for a stronger analysis and a weaker presentation. It is one scenario, not a universal office benchmark. GDPval-AA moved from 1,605 to 1,695 Elo on the high-versus-xhigh pair, and xAI's GDPval chart puts Fable 5.1 max at 1,735.

Artificial Analysis post on AA-Briefcase scores and example deck costs for Grok 4.7 and Grok 4.6
Artificial Analysis on X, September 22, 2026: analytical Elo up, presentation Elo down, example decks about $8 versus $4.40 at xhigh.

A lower hallucination rate, not a higher accuracy rate

AA-Omniscience hallucination rate was 29% for Grok 4.7 xhigh and 34% for Grok 4.6 high. Accuracy was 47% versus 48%. The index moved from 30 to 32. If you need fewer confident wrong answers, that is a real shift. If you need more correct answers, this pair does not show it. The effort labels still differ.

Where Grok 4.6 is still the better default

Short work, where the extra thinking is the product

Paying the same rate for about twice the output is a good trade only when the extra output prevents a retry. For a bounded extraction, a short patch, or a status JSON, Grok 4.6's shorter trace is the feature. Hacker News readers looking at the intelligence-token chart called the change a regression in token efficiency. That comment does not add a new measurement. It is the same Artificial Analysis chart, read as a cost problem rather than a capability upgrade.

Hacker News comment on Grok 4.7 token efficiency versus Grok 4.6
notduckrabbit on Hacker News: the intelligence-index token chart reads as a regression against Grok 4.6.

On the tasks Artificial Analysis called out as regressions versus Grok 4.6 high, AA-LCR and AutomationBench-AA, there is no reason to leave 4.6. Absolute scores were not published, so the size of those drops is the published delta only: 3.7 and 1.1 points.

Cursor plans, if Fast and effort are not pinned

A list rate of $2/$6 does not describe a Cursor Ultra percentage. Dynamix86's timings are a single-user log: Grok 4.7 intervals of 34, 55, and 38 minutes per percentage point, against 133 and 101 minutes on Grok 4.6 earlier the same day, both at extra high with Fast off. The author also read an Artificial Analysis chart as $8.82 versus $3.53 per coding-agent task. Those dollar figures were the author's reading of a chart. They were not printed as a sentence in the Artificial Analysis article opened for this page, so they stay attributed to the post.

A reply on the same thread was blunter: the release looked poor because the cost was clearly too high even if benchmarks were set aside.

Reddit comment by truecakesnake saying Grok 4.7 costs too much
u/truecakesnake on r/cursor: the objection is the bill, not a named benchmark.

Cursor's docs also say Fast is the default on Pro and higher, at double the token rates. If one side of a personal test is Fast and the other is not, the 2× price is a setting, not a model change. The Reddit log says Fast was off. Many casual comparisons will not.

Anything you already accept from Grok 4.6

Grok 4.6's review notes already describe a 500K agent model at this rate. Grok 4.7 does not add a new context tier on the API card, and it does not cut the list price. If 4.6 already finishes the job and you are not blocked on long terminal work or long documents, the upgrade is optional. HealthBench Professional still trails GPT-5.6 Sol and Fable 5.1 on xAI's own table (56.7 versus 60.5 and 62.1), so a clinical-reasoning switch is not the story either.

What people actually said

The early comments split into a cost camp and a feel camp. Both are hours old. Neither camp published a prompt, a repository, or a score.

The bill. Dynamix86's 2.5× plan-usage estimate and truecakesnake's "cost is clearly too high" sit next to the token chart. A separate Hacker News commenter, notduckrabbit, pointed at the same efficiency gap. Together they say: do not migrate a high-volume route on launch day.

The feel. Other r/cursor comments, posted the same afternoon, liked the model and did not mention price.

Reddit comment saying Grok 4.7 is not as slow as Grok 4.6
u/vandersky_ : an early speed impression, with no timing log.
Reddit comment comparing Grok 4.7 to a less overthinking Opus
u/kodka: a feel comparison with Opus, not a shared task.
Reddit comment preferring Devin SWE 2.0 over Grok 4.7
u/ImmediateAttention88: prefers Devin's SWE 2.0 with Fusion, and is still using both agent IDEs.

The speed and "does not overthink" comments can both be true for a chat session and still false for an xhigh coding agent that emits 81k tokens. "Not as slow as 4.6" was not paired with a token count. Artificial Analysis measured about 188 tokens per second on long prompts and about 7.1 minutes of decode time per intelligence task, excluding time to first token and overhead. A fast first impression and a long agent trace are compatible.

YouTube search for "Grok 4.7 vs Grok 4.6" on September 22 returned launch explainers, including pre-release videos that guess parameter counts. Those counts are not on the xAI launch page, which only says "a new, larger base model." No video comment is used here.

The verdict

There is no overall winner. Pick by the shape of the work, and keep the effort label in the decision.

WorkloadChooseWhyWhat to watch
Short extraction, classification, or a small patchGrok 4.6The index gain is small and the output trace is much shorterDo not pay xhigh for a one-line answer
Repeated Cursor work on a usage poolGrok 4.6, unless a task is failingOne Ultra log showed about 2.5× plan usage at extra high with Fast offPin Fast and effort before blaming the model
Multi-hour coding in Grok Build or a similar agentPilot Grok 4.7 xhighCoding Agent Index 56 vs 47 at matched xhighCompare completed tasks, not the index alone
Market models, memos, and target decksPilot Grok 4.7, then edit the deckAnalytical Elo jumped; presentation Elo slipped; example decks were about $8 vs $4.40Budget a human pass on slides
Long-context API calls above 200KEither, on priceBoth cards double the rates past 200KCursor's 256K boundary is a different meter
You need Fable-level CursorBench or Terminal-BenchNot this pairFable 5.1 Max is 51.8% and 57.9% on xAI's tablePrice that choice on Fable's card, not Grok's

A practical test is one accepted task, same effort, Fast off, same harness. Record completed or not, human edits, output tokens, and either API dollars or plan percentage. One success is not a rate.

Run the same task in Tabbit before you migrate a route

A benchmark row does not tell you how either model behaves on your tabs, docs, and half-finished research. Tabbit Browser is the workspace for that check: the model can sit next to the pages you are already reading, instead of in a separate chat that never sees them. The agentic browser guide and the agentic reasoning guide separate a model score from a finished multi-step workflow. The AI browser comparison and the Tabbit browser guide cover the product around the model. The 2026 AI browser roundup is the wider set if Tabbit is not the client you want.

Tabbit new chat showing several models answering what is Tabbit side by side
Tabbit can put more than one model on the same question. This screenshot shows other models, not a Grok 4.7 versus Grok 4.6 run.

Both Grok IDs are registered for Tabbit. Whether Grok 4.7 or Grok 4.6 appears in your picker depends on the account and the current selector. This article does not include a Tabbit transcript, a token log, or a winner. If both IDs are available, run one task you already know how to grade, keep effort and tools the same, and keep Grok 4.6 if the only change is a longer trace. The Tabbit practices guide is about the browser workflow, not a claim that either Grok build is preinstalled for every account.

API list price, a Cursor pool, and Tabbit access are three different bills. Do not add them together.

Tabbit Browser

FAQ

Is Grok 4.7 better than Grok 4.6?

It depends on the task and the reasoning effort. On Artificial Analysis's Coding Agent Index, Grok 4.7 xhigh with Grok Build scored 56 versus 47 for Grok 4.6 xhigh. The Intelligence Index only moved from 44 at Grok 4.6 high to 46 at Grok 4.7 xhigh, and some non-agent tasks regressed. There is no single winner.

Do Grok 4.7 and Grok 4.6 cost the same?

On the xAI model cards opened September 22, 2026, both list $2 per million input tokens, $0.50 cached input, and $6 output below 200K context, then $4, $1, and $12 above 200K. Cursor documents a separate usage pool and a Fast tier. The list rate matching does not mean a finished task costs the same.

Why can Grok 4.7 cost more if the token price is unchanged?

Artificial Analysis measured about 81,000 output tokens per Intelligence Index task for Grok 4.7 xhigh, versus about 36,000 for Grok 4.6 high and about 38,000 for Grok 4.6 xhigh. At $6 per million output tokens, that output alone is roughly $0.49 versus $0.22 or $0.23, before input, cache, tools, or retries.

Which Grok 4.7 benchmark gains are on the same effort level?

The Coding Agent Index comparison is xhigh versus xhigh: 56 versus 47, with DeepSWE 73% versus 65%, Terminal-Bench 4.0 33% versus 18%, and SWE-Atlas-QnA 63% versus 58% inside Grok Build. xAI's launch table compares Grok 4.7 xHigh with Grok 4.6 High, and DeepSWE on that table is marked high effort for Grok 4.7. Do not treat those rows as a same-rung test.

Should I switch from Grok 4.6 to Grok 4.7 in Cursor?

Switch a long coding or document task if you can accept a larger usage drop. One Cursor Ultra user, on extra high and with Fast off, estimated about 2.5 times the plan usage and planned to move most work back to 4.6. Other early comments said 4.7 felt faster or less prone to overthinking. Pin the effort and Fast setting before you compare.

Can I compare Grok 4.7 and Grok 4.6 in Tabbit Browser?

Both models have Tabbit model pages. Whether they appear in your selector depends on the account and the live picker. This article does not include a same-task Tabbit run, so a selector entry is not evidence that one model wins a workflow.

Take the next step

Let Tabbit work alongside you.

Research across tabs, automate repetitive browser work, and keep every piece of context within reach.