BenchLM overall public score: 63.4 / 100.
Public leaderboard rank: #43 / 218.
Agentic category: #7 / 130.
Coding category: #18 / 135.
Public evidence rows: 42; of these, Agentic 4/4 and Coding 12/12 are marked as verified.
Speed: approximately 66 tokens/s, with a time to first token of approximately 32.3 seconds.
Context: 500K tokens.
Pricing: $2 input, $6 output, and $0.50 / 1M tokens for cached input.
The Coding results listed on the page include DeepSWE at 65.9%, CursorBench 3.2 at 70.8%, and FrontierCode 1.1 Extended at 61.3%. GPT-5.6 Sol is the best comparison on DeepSWE at 72.7%, while the CursorBench 3.2 page marks Grok 4.6 as the current best verified value.
BenchLM's value is that it puts model scores, pricing, speed, context, and evidence coverage on the same page. It also explicitly warns that there is currently insufficient data for Grok 4.6 in categories such as Reasoning, Knowledge, and Math, so the overall score should not be understood as a complete profile of all its capabilities.
Grok 4.6