In Cursor, the author used Grok 4.6 Extra High and GPT-5.6 Sol Medium to execute the same detailed backend plan, starting from the same point and using the same plan, with a task size of approximately 2,500 lines of code; Fable 5 High served as an independent reviewer.
The author's approximate scores were Sol 60, Grok 40. Sol performed better on money-related edge cases, race-condition risks, and overall architecture, and its testing was more targeted.
In terms of cost, the author said Grok's Cursor usage barely changed, while Sol used about 5% of the $200 monthly subscription allowance. Commenters cautioned that run order, leftover branches, and context contamination could affect the result; the author replied that Grok ran first and that they had quickly checked Sol's reasoning process, finding no evidence that it had read Git history.
Some users felt that if Grok 4.6 failed on its first attempt, its low cost and speed would make a second attempt acceptable.
Some users pointed out that Grok 4.6 may consume more reasoning tokens and tool calls than 4.5, so "the same price per token" does not mean "the same cost per task."
The discussion also noted that different harnesses, such as Cursor and Codex, can change model performance.
This is a small-sample engineering experience based on a single task, and is insufficient to overturn public leaderboards. But it clearly shows Grok 4.6's boundary: in complex backend implementation, its cost advantage is significant; on money logic, race conditions, and architecture-level edge handling, Sol may be more reliable.
Grok 4.6