The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73%.
The post also discusses version changes in Terminal-Bench: after the move from 2.1 to 3.0, models that had originally scored close to 90% could fall back to around 30%. Scores from different versions therefore cannot be compared directly.
One user uses Grok 4.6 in Cursor Enterprise and considers it suitable for work involving large numbers of tokens.
One user considers Grok 4.6's capabilities close to Opus 5's, but says the 500K context window remains limiting for some architect-type work.
Another group of comments uses Grok 4.6 for simple E2E tasks and sub-Agents, describing it as "smart enough and very fast."
Some comments argue that DeepSWE reflects only repository tasks and hidden tests, and cannot represent architecture, maintainability, or long-term collaboration in real development.
This post is more of a quick community interpretation of public scores than an independent rerun. Its value lies in reminding readers to consider DeepSWE, the Terminal-Bench version, and the test objective together when assessing Grok 4.6's coding results, and to evaluate "speed suitable for sub-Agents" separately from "stability when completing complex engineering work independently."
Grok 4.6