The post primarily discusses real-world impressions in light of xAI/Artificial Analysis scorecards. A highly engaged comment describes Grok 4.6 as "cheap and fast" for coding and shares a workflow in which Opus handles planning while Grok handles specific, narrowly scoped changes.
Another commenter reported the opposite experience: after running Grok 4.6 and GPT-5.6 Sol in parallel in Cursor, they felt there was still a clear gap between the two. They said they had used roughly 200 million tokens with Grok, 300 million with Sol, and 300 million with Opus. This feedback does not disclose a publicly reproducible task set and should be treated as a long-term user observation rather than a formal benchmark.
The post cites Artificial Analysis's "cost per task" view rather than comparing only the price per million tokens.
A commenter says Grok 4.6's per-task cost is roughly comparable to Kimi K3's; another comment suggests it may even be lower.
The community broadly notes that whether a model is "verbose," the number of retries, cache hits, and the number of tool calls can all change the actual bill.
Developers' main perception of Grok 4.6's advantages centers on speed, price, and coding usability.
Community experiences are inconsistent: in Cursor, some users consider it close to frontier models, while others find Sol more reliable at handling edge cases.
Real-world selection requires tracking task success rate, retry count, tool calls, elapsed time, and total cost, rather than looking only at leaderboard rank.
Grok 4.6