Grok 4.6 · Community source · Personal experience
This evidence note covers “Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
In Cursor, the author used Grok 4.6 Extra High and GPT-5.6 Sol Medium to execute the same detailed backend plan, starting from the same point and using the same plan, with a task size of approximately 2,500 lines of code; Fable 5 High served as an independent reviewer.
The author's approximate scores were Sol 60, Grok 40. Sol performed better on money-related edge cases, race-condition risks, and overall architecture, and its testing was more targeted.
In terms of cost, the author said Grok's Cursor usage barely changed, while Sol used about 5% of the $200 monthly subscription allowance. Commenters cautioned that run order, leftover branches, and context contamination could affect the result; the author replied that Grok ran first and that they had quickly checked Sol's reasoning process, finding no evidence that it had read Git history.
Some users felt that if Grok 4.6 failed on its first attempt, its low cost and speed would make a second attempt acceptable.
Some users pointed out that Grok 4.6 may consume more reasoning tokens and tool calls than 4.5, so "the same price per token" does not mean "the same cost per task."
The discussion also noted that different harnesses, such as Cursor and Codex, can change model performance.
This is a small-sample engineering experience based on a single task, and is insufficient to overturn public leaderboards. But it clearly shows Grok 4.6's boundary: in complex backend implementation, its cost advantage is significant; on money logic, race conditions, and architecture-level edge handling, Sol may be more reliable.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Reddit r/cursor · u/Rashe39 · Original publication date Unknown · Site edit date 2026-09-20
Open original sourceGrok 4.6
Download the Tabbit client to check model access