Grok 4.6 · Media / benchmark · Independent measurement
This evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
BenchLM overall public score: 63.4 / 100.
Public leaderboard rank: #43 / 218.
Agentic category: #7 / 130.
Coding category: #18 / 135.
Public evidence rows: 42; of these, Agentic 4/4 and Coding 12/12 are marked as verified.
Speed: approximately 66 tokens/s, with a time to first token of approximately 32.3 seconds.
Context: 500K tokens.
Pricing: $2 input, $6 output, and $0.50 / 1M tokens for cached input.
The Coding results listed on the page include DeepSWE at 65.9%, CursorBench 3.2 at 70.8%, and FrontierCode 1.1 Extended at 61.3%. GPT-5.6 Sol is the best comparison on DeepSWE at 72.7%, while the CursorBench 3.2 page marks Grok 4.6 as the current best verified value.
BenchLM's value is that it puts model scores, pricing, speed, context, and evidence coverage on the same page. It also explicitly warns that there is currently insufficient data for Grok 4.6 in categories such as Reasoning, Knowledge, and Math, so the overall score should not be understood as a complete profile of all its capabilities.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
BenchLM.ai · Author not disclosed · Original publication date Unknown · Site edit date 2026-09-20
Open original sourceGrok 4.6
Download the Tabbit client to check model access