Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Grok 4.6 · Media / benchmark · Independent measurement

BenchLM: Grok 4.6's Public Scores, Speed, and Cost

This evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Test conditions
The BenchLM snapshot records 63.4/100, about 66 tokens/s, and roughly 32.3 seconds to first token; evidence is insufficient for several categories including Reasoning, Knowledge, and Math.
Source boundary
Supports reading speed, TTFT, and the composite only within the page’s evidence coverage, and helps define retest questions.
Unsupported claims
Does not support treating the composite as a complete capability profile or making claims about later prices or unmeasured categories.

Key data and applicable tasks

Key results

  • BenchLM overall public score: 63.4 / 100.

  • Public leaderboard rank: #43 / 218.

  • Agentic category: #7 / 130.

  • Coding category: #18 / 135.

  • Public evidence rows: 42; of these, Agentic 4/4 and Coding 12/12 are marked as verified.

  • Speed: approximately 66 tokens/s, with a time to first token of approximately 32.3 seconds.

  • Context: 500K tokens.

  • Pricing: $2 input, $6 output, and $0.50 / 1M tokens for cached input.

Specific benchmarks

The Coding results listed on the page include DeepSWE at 65.9%, CursorBench 3.2 at 70.8%, and FrontierCode 1.1 Extended at 61.3%. GPT-5.6 Sol is the best comparison on DeepSWE at 72.7%, while the CursorBench 3.2 page marks Grok 4.6 as the current best verified value.

Interpretation

BenchLM's value is that it puts model scores, pricing, speed, context, and evidence coverage on the same page. It also explicitly warns that there is currently insufficient data for Grok 4.6 in categories such as Reasoning, Knowledge, and Math, so the overall score should not be understood as a complete profile of all its capabilities.

What this supports

  • Supports reading speed, TTFT, and the composite only within the page’s evidence coverage, and helps define retest questions.

What this does not support

  • Does not support treating the composite as a complete capability profile or making claims about later prices or unmeasured categories.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

BenchLM.ai · Author not disclosed · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Grok 4.6

Compare Grok 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Grok 4.6: What Changed, What It Costs, and Who It Fits

A sourced guide to Grok 4.6: the 500K context model, benchmark-version split, live API price and a safer pilot decision.

Comparison · English

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

Related reviews

Reddit r/singularity: Grok 4.6 Benchmarks and Real-World Coding FeedbackKey points from the post and comments The post primarily discusses real-world impressions in light of xAI/Artificial Analysis scorecards. A highly engaged comment describes Grok 4.6 as "cheap and fast" for coding and shares a workflow in which Opus handles pla。X: Same-Prompt Cost and Speed Comparison of Grok 4.6 and Opus 5The author says they tested Grok 4.6 and Opus 5 with the same prompt: Grok 4.6: approximately $1.74, completed in about 5 minutes.。Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Grok 4.6 Official Release: Benchmarks and Capability EvaluationThis evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Build Fast with AI: General Methods from 100 Grok PromptsThe article's five rules Use Grok's real-time search and ask it to provide sources. When you need public-opinion, trend, or public-reaction data, explicitly ask it to search X.。Reddit r/LoveGrok: Practical Project Instructions and Positive ConstraintsCommunity experience One user suggested putting custom instructions in Project instructions rather than ordinary settings, because project instructions have weighting and context better suited to long-term writing.。X: Anshu's Grok 4.6 App and Design Workflow with the Same PromptThe author ran the same one-shot prompt in Grok Build with Grok 4.5 and Grok 4.6, then compared the resulting apps and designs. The author considers 4.6 a clear improvement over 4.5, saying it can even alternate for the lead with Fable on design tasks.。X: Eric Zakariasson's Short Prompts and Strict Verification for Grok 4.6The author compares long and short prompts, as well as wording such as “work very hard.” The conclusion is that the specific wording itself has little impact. Long prompts can add specificity and suit tasks where the user already knows what they want, while Gr。