TabbitBlog

Grok 4.7 Review: Same $2/$6 Price, About Twice the Tokens

A Grok 4.7 review of the unchanged $2/$6 rates, the jump to about 81k output tokens, and which workloads justify the extra work.

In this article
  1. The one number that decides this review
  2. Key takeaways
  3. The benchmarks, and how much to trust them
  4. Where it is genuinely good
  5. Analysis got better, and the slides did not
  6. Grok Build's coding agent moved more than the index
  7. The rate card is still the cheap flagship cell
  8. Where it bites
  9. The token bill, not the rate card
  10. Most hard terminal tasks still fail
  11. The index step is small, and the effort mismatch is easy to miss
  12. What people actually said
  13. The bill
  14. The feel
  15. Try the workflow before you trust the row
  16. Choose by workload

Grok 4.7 is xAI's September 21, 2026 coding and knowledge-work model. The useful question is not whether the launch chart has a winning row. It is whether the model is worth leaving Grok 4.6 for, once the bill includes how much it writes.

On Hacker News, dumberquestions put the trap in one line: "Token price doesn't tell you much without knowing token efficiency." That is the review. The previous generation is in the Grok 4.6 article. A pricing article and an alternatives list are not on this site yet. Nothing below is a controlled run inside Tabbit Browser.

The one number that decides this review

Two figures point in opposite directions, and both were visible on September 22, 2026.

The launch page prices Grok 4.7 xHigh at $2 per million input tokens and $6 per million output tokens, the same cells as Grok 4.6 High. The API docs repeat $2.00 and $6.00 for grok-4.7.

Artificial Analysis then measured Grok 4.7 at xhigh using about 81,000 output tokens per Intelligence Index task. Grok 4.6 at xhigh used about 38,000. The same write-up puts Grok 4.6 high near 36,000 and GPT-6 Astra max near 27,000. The sticker did not move. The amount of writing did.

The index only moved from 44 for Grok 4.6 high to 46 for Grok 4.7 xhigh. The Coding Agent Index, run in Grok Build, moved more: 56 versus 47. If your task is a short answer, you are paying for thinking you did not order. If the last model stalled halfway through a repository, the extra tokens are the product.

Key takeaways

  • The public list price matches Grok 4.6. The xhigh token count on the Intelligence Index does not.

  • The clearest gain is agentic: Grok Build's coding-agent score rose 9 points, and office-work analysis quality rose sharply.

  • The general index is a 2-point step and still trails Claude Fable 5.1, GPT-6 Astra, and GPT-5.6 Sol on that chart.

  • Slide quality went the other way. Presentation Elo fell while analytical Elo jumped.

  • Cursor plan burn and API dollars are different meters. One Ultra user saw faster plan drain. That does not set the API price.

The benchmarks, and how much to trust them

Read each row with its harness. A Terminal-Bench percent from xAI's table is not the Terminal-Bench percent from Grok Build.

SignalGrok 4.7The comparison that mattersCondition, checked 2026-09-22
List price$2 / $6 per 1MSame as Grok 4.6. Sol Max $4 / $20. Fable Max $10 / $50xAI launch table, xHigh column
API modelgrok-4.7Context 500,000. Default effort is high. xhigh existsdocs.x.ai at a glance
Intelligence Index46Grok 4.6 high 44. Sol max 47. Fable 5.1 and Astra 53Artificial Analysis, xhigh
Output tokens per index taskabout 81kGrok 4.6 xhigh about 38k. Astra max about 27kSame article. Not a dollar invoice
Coding Agent Index56Grok 4.6 xhigh in Grok Build 47. Fable and Astra 62Native harness, not the index harness
AA-Briefcase Elo1,657Grok 4.6 high 1,546. Fable 5.1 1,678xhigh versus high on the 4.6 side
Analysis vs slides1,994 / 1,499 EloGrok 4.6 high 1,690 / 1,519Analysis up, presentation down
CursorBench 4.046.3%, about $6.01 / taskFable 5.1 Max 51.8%, about $17.28xAI scatter, Extra High. Vendor harness
Terminal-Bench 4.038.0%Grok 4.6 High 20.3%. Sol Max 37.3%. Fable Max 57.9%xAI table, xHigh column
Terminal-Bench in Grok Build33%Grok 4.6 xhigh 18%Artificial Analysis. Do not average with 38%

Three limits sit on top of that table.

The effort labels do not match. Several "Grok 4.7 beat Grok 4.6" lines compare xhigh with high. The token double, 81k versus 38k, is the cleaner pair because both sides are xhigh. The +2 index point is not that pair.

The harness changes the story. Coding Agent numbers include Grok Build. The Intelligence Index uses a shared harness. xAI's Terminal-Bench cell is 38.0 percent. Artificial Analysis, inside Grok Build, prints 33 percent against an 18 percent baseline. The Decoder's 26 percent is a third caption and is not used here.

Vendor charts pick their rivals. CursorBench's scatter, as read on the launch page, shows Fable, Opus, Sol, and Sonnet. It does not show GPT-6 Astra. A low cost-per-task on that chart is not a win over Astra. Benchmarks are a direction, not a ruler.

There is no Tabbit task log to put above this table. The next sections stay with public measurements and named comments.

Where it is genuinely good

Analysis got better, and the slides did not

The surprise is not the 46 on the index. On AA-Briefcase, analytical quality moved from 1,690 Elo for Grok 4.6 high to 1,994 for Grok 4.7, while presentation quality moved from 1,519 to 1,499. Artificial Analysis said the same split in a September 22 post: analytical Elo 1,698 to 1,994 on their Lite scenario, presentation 1,531 to 1,499. The two write-ups do not use the same 4.6 baseline. Both say the writing of the analysis improved and the deck finish did not.

Artificial Analysis post on X showing Grok 4.7 AA-Briefcase Elo of 1657, a gain in analytical quality, and a drop in presentation quality
Artificial Analysis on X, September 22, 2026. The chart is their Elo snapshot. The 1698 and 1531 baselines in the post text are not the same pair as the article's 1690 and 1519.

Original post.

If the deliverable is a market model, a memo, or a spreadsheet argument, that split matters more than a two-point index. If the deliverable is a polished deck, Grok 4.7 is the wrong thing to celebrate.

Grok Build's coding agent moved more than the index

Inside Grok Build, Artificial Analysis scored DeepSWE v1.1 at 73 percent versus 65 percent, Terminal-Bench 4.0 at 33 percent versus 18 percent, and SWE-Atlas-QnA at 63 percent versus 58 percent, all against Grok 4.6 xhigh. The composite is 56, fourth among the native harnesses they plotted, behind Fable 5.1 and GPT-6 Astra.

That is a real step for long coding loops. It is still a system score: model plus Grok Build. It does not say a raw grok-4.7 API call, on high rather than xhigh, will clear the same tasks. For how those loops differ from a single answer, the agentic reasoning note is the adjacent explainer.

The rate card is still the cheap flagship cell

On xAI's CursorBench scatter, Grok 4.7 Extra High sits at 46.3 percent and about $6.01 per task. Fable 5.1 Max sits at 51.8 percent and about $17.28. Grok 4.7 High is 43.9 percent at about $4.69, and Low is 33.1 percent at about $1.58. You can buy a lower score for a lot less money than Fable's top setting, or you can buy Fable's higher score. Those are different purchases. The Fable 5.1 review and the GPT-6 Astra review cover the expensive end of that choice.

Where it bites

The token bill, not the rate card

About 81,000 output tokens at $6 per million is roughly $0.49 of output before input, cache, tools, or retries. About 38,000 is roughly $0.23. That arithmetic uses only the output counts Artificial Analysis printed. It is not your invoice.

A Cursor Ultra subscriber, Dynamix86, ran what they called the same task on extra high with fast mode off and concluded Grok 4.7 burned plan percentage about 2.5 times as fast as Grok 4.6. They also read a chart as $8.82 versus $3.53 per task. Treat the plan timing as their log. Treat the dollar pair as their reading of a picture, not as a figure this review remeasured.

Reddit post by Dynamix86 saying Grok 4.7 used Cursor Ultra quota about 2.5 times as fast as Grok 4.6 on extra high
Dynamix86 on r/cursor, September 21, 2026. Extra high, fast mode off, one account. The dollar figures in the post are the author's reading of a chart.

Original thread.

xAI also sells a fast variant at twice the token rates, and only in Cursor and Grok Build. The public API does not offer it. The US regional endpoint adds a 10 percent usage premium. Stacking "fast" and a long prompt is a different product from the $2 / $6 headline.

Most hard terminal tasks still fail

Official Terminal-Bench 4.0 at xHigh is 38.0 percent. Fable 5.1 Max on that same table is 57.9 percent. Sol Max is 37.3 percent. Grok 4.7 improved on Grok 4.6 High's 20.3 percent and still misses most of the set. HealthBench Professional is 56.7 percent against Fable's 62.1 percent and Sol's 60.5 percent. Harvey's legal-agent cell is 19.6 percent, far above Sol's 2.5 percent and still a low absolute.

A model trained to stay on multi-hour work can remain a model that fails the hour. Do not read "better than 4.6" as "finished the job."

The index step is small, and the effort mismatch is easy to miss

Forty-six versus forty-four will not change a short chat, a translation, or a one-file edit. GPT-5.6 Sol max is already at 47 on that index, with a higher list price. Teams who wanted a general leap will not find it in this release. Teams who compare xhigh against high, then announce a broad win, are reading the label they prefer.

Hallucination rate improved, 29 percent versus 34 percent for Grok 4.6 high, while accuracy stayed near 47 percent versus 48 percent. Fewer wrong answers among non-correct replies is not the same as knowing more.

What people actually said

The comments split into a bill argument and a feel argument. They do not settle the benchmarks.

The bill

Hacker News comments saying token price is not token efficiency, and that Grok 4.7's cost per task is a tough sell against Fable 5.1 Low
dumberquestions and user43928 on the Grok 4.7 thread, September 22, 2026.

dumberquestions and user43928.

saejox was blunter about a model that costs more per token and still uses fewer of them: "Not even close to astra. It is expensive, but uses way fewer tokens."

Hacker News comment by saejox saying Astra uses fewer tokens than Grok 4.7 despite a higher price
saejox on Hacker News. No task log is attached.

Comment.

The feel

rayiner likes Grok, as opposed to code, for legal research because it gets to the point faster than Opus 5. The sentence is about the recent Grok line, not a scored 4.7 case file.

Hacker News comment by rayiner preferring Grok for legal research because it gets to the point
rayiner on Hacker News. A preference, not a legal-accuracy test.

Comment.

The first-hour thread on r/cursor is the same split in miniature. vandersky_ wrote "Not as fucking slow as 4.6." kodka wrote that it feels like Opus 5 without the overthinking. exploriosapp is still testing and thinks the writing is better. ImmediateAttention88 prefers Devin's SWE 2.0 with a fusion API and is using both tools. Nobody posted the repository, the effort level, or a timer.

r/cursor comments on the first hours of Grok 4.7, mixing speed praise, writing notes, and a preference for Devin
r/cursor, September 21, 2026. Early impressions with no shared task.

Thread.

The editorial read matches the numbers. People who watch the meter are unhappy. People who wanted a less ponderous coding feel are curious. Neither group has posted a repeated task score.

Try the workflow before you trust the row

A benchmark row will not show how Grok 4.7 behaves on the pages already open in your browser. Tabbit Browser is the client for that check: the model picker, the live page, and the file sit in one place. The AI browser guide and the Tabbit Browser walkthrough describe that layout. Agentic browsing and browser automation are the neighboring jobs, and the 2026 AI browser shortlist is there if the question is the browser rather than the model.

The boundary is simple. This review did not run a fixed extraction, a repeated comparison, or a timed page task in Tabbit. If Grok 4.7 is absent from your picker, the button below does not create access. API dollars, a Cursor plan, and a Tabbit account are three prices. A single lucky answer would not become a success rate even if it had been captured.

Choose by workload

WorkloadUse Grok 4.7?WhyWatch
Long coding loop that 4.6 abandonsYes, in Grok Build or Cursor, start below xhighThe coding-agent step is the largest public gainPlan percentage and output tokens
One-file edit or chatNoThe index moved 2 points and the token count jumpedStay on 4.6 or a smaller model such as the one in the Gemini 3.8 Flash review
Memo, model, or spreadsheet argumentYes, if analysis is the productAnalytical Elo jumpedDo not expect a better slide finish
Terminal suite where Fable already clears the runNo, unless price dominates38% official Terminal-Bench still fails most tasksDo not mix in the 33% Grok Build figure
Legal research where brevity mattersTry, then check sourcesOne practitioner prefers the tone. Harvey's agent cell is only 19.6%Tone is not correctness
Work that lives in open tabsTry in Tabbit only if the picker lists itThe client is the missing layer, not a higher scoreNo success rate is claimed here

Grok 4.7 is the upgrade when the previous model stops, and you can pay for the extra tokens. It is an expensive habit when the previous model already finished.

Tabbit Browser

FAQ

Is Grok 4.7 worth switching to from Grok 4.6?

Switch when a long coding-agent run or an analytical document fails on Grok 4.6 and you can afford more output tokens. The list price is still $2 per million input tokens and $6 per million output tokens. At xhigh, Artificial Analysis measured about 81,000 output tokens per Intelligence Index task, against about 38,000 for Grok 4.6 xhigh. Routine edits do not need that extra work.

Why do people say Grok 4.7 costs more if the token price did not change?

The rate card and the bill measure different things. xAI lists the same $2 and $6 rates as Grok 4.6. Independent testing shows Grok 4.7 xhigh writes about twice as many output tokens on the same index. A Cursor Ultra user on extra high, with fast mode off, also watched plan percentage fall faster. That usage note is one account, not an API invoice.

How does Grok 4.7 compare with Claude Fable 5.1 and GPT-6 Astra?

On the Artificial Analysis Intelligence Index, Grok 4.7 xhigh scores 46. Claude Fable 5.1 and GPT-6 Astra score 53, and GPT-5.6 Sol scores 47. Grok Build plus Grok 4.7 scores 56 on the Coding Agent Index, behind Fable and Astra at 62. The gap is smaller on some office-work Elo charts than on the general index.

What is Grok 4.7 Fast, and can the API use it?

xAI's docs describe Grok 4.7 Fast as the same model on faster infrastructure, billed at twice the standard token rates. It is available in Cursor and Grok Build, it is not part of Grok Build's free quota, and it is not on the public xAI API. The public model name is grok-4.7.

Does a higher CursorBench score mean Grok 4.7 wins on coding?

No. xAI's own CursorBench chart puts Grok 4.7 Extra High at 46.3 percent and about $6.01 per task, below Fable 5.1 Max at 51.8 percent and about $17.28. Official Terminal-Bench 4.0 is 38.0 percent, so most of those terminal tasks still fail. A second Terminal-Bench number, 33 percent, comes from Grok Build and should not be mixed into the same cell.

Can I run Grok 4.7 in Tabbit Browser?

Tabbit's Grok 4.7 page says you can use the model when it appears in your account's model picker. This review did not run a fixed task in Tabbit, so it does not report a success rate, latency, or client cost. API rates and a Cursor plan are separate from whatever your Tabbit account includes.

Take the next step

Let Tabbit work alongside you.

Research across tabs, automate repetitive browser work, and keep every piece of context within reach.