Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Grok 4.6 · Media / benchmark · Independent measurement

Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6

This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Test conditions
Artificial Analysis records Intelligence Index 61, Terminal-Bench v2.1 at 88.4%, and about $0.84 per task; full sample, reasoning tier, and private AA-Briefcase harness details are not public.
Source boundary
Supports a bounded intelligence/cost reading of this AA snapshot; it does not support merging v2.1 with xAI v3.0 or other harness scores.
Unsupported claims
Does not support current pricing, Tabbit availability, a general success rate, or production-safety guarantees.

Key data and applicable tasks

Results at a glance

Artificial Analysis rates Grok 4.6 at 61 on its Intelligence Index, tied with GPT-5.6 Sol and behind only Claude Opus 5 and Fable 5. The firm's interpretation is that Grok 4.6's advantage is not limited to static reasoning; it lies in its overall performance and cost efficiency in Agent scenarios such as knowledge work, customer-service tool calls, and terminal tasks.

Key data

  • Artificial Analysis Intelligence Index: 61, 5 points higher than Grok 4.5 and 23 points higher than Grok 4.3.

  • GDPval-AA v2: 1,753 Elo, second only to Claude Opus 5; its confidence interval overlaps with those of Fable 5 and Qwen3.8 Max.

  • τ³-Banking: 50.7%, in the highest-scoring group and close to Qwen3.8 Max at 51.3%.

  • Terminal-Bench v2.1: 88.4%, at the same level as the leading models.

  • AA-Briefcase: 1,577 Elo, close to Fable 5-level long-horizon knowledge-work capability.

  • Cost per task: approximately $0.84.

  • AA-Briefcase averaged about 53 turns and 0.5B input tokens; the comparison Claude Opus 5 Max used about 103 turns and 2.0B input tokens.

Cost context

Grok 4.6 is listed at $2 / 1M input tokens and $6 / 1M output tokens, with cached input at $0.5 / 1M tokens. Artificial Analysis considers Grok 4.6 more advantageous on output pricing and task cost than GPT-5.6 Sol at $5 / $30 and Claude Opus 5 at $5 / $25.

Interpretation of the evaluation

Artificial Analysis's results support the following usage judgments:

  1. Grok 4.6's advantages are more apparent when a task requires multiple rounds of retrieval, source organization, customer-service tool calls, and long-chain knowledge work.

  2. Do not look only at the list price per million tokens; also consider the number of turns and tokens needed to complete the same task.

  3. Terminal-Bench v2.1 and the v3.0 version released by xAI are different versions, so 88.4% and 26% should not be treated as contradictory figures.

Notes

AA-Briefcase is a private evaluation by Artificial Analysis and cannot be treated as a public benchmark. All results should still be retested against your own codebase, toolchain, and long-task samples.

What this supports

  • Supports a bounded intelligence/cost reading of this AA snapshot; it does not support merging v2.1 with xAI v3.0 or other harness scores.

What this does not support

  • Does not support current pricing, Tabbit availability, a general success rate, or production-safety guarantees.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Author not disclosed · Original publication date 2026-08-12 · Site edit date 2026-09-20

Open original source

Grok 4.6

Compare Grok 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Grok 4.6: What Changed, What It Costs, and Who It Fits

A sourced guide to Grok 4.6: the 500K context model, benchmark-version split, live API price and a safer pilot decision.

Comparison · English

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

Related reviews

KIE: Grok 4.6 Release Evaluation and Capability BreakdownArticle conclusion KIE summarizes Grok 4.6's Intelligence Index as 61, tied with GPT-5.6 Sol; its price is $2 for input and $6 for output per 1M tokens. The article argues that its main selling point is its price-to-intelligence ratio, rather than leading on e。BenchLM: Grok 4.6's Public Scores, Speed, and CostThis evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Grok 4.6 Official Release: Benchmarks and Capability EvaluationThis evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Emergent: Breaking Down Grok 4.6's Evaluation ResultsEmergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61, tied with GPT-5.6 Sol; but on DeepSWE 。X: Nikolai Yakovenko on the Single-Prompt Playable-Game PatternThe author believes that making a playable game with a “one-sentence prompt” has become a common benchmark when a new model or Agent harness is released. The comments showed a case in which Grok 4.6 generated a simulation of Phantasy Star 1 and an AI Agent ope。Venice: Four Grok 4.6 Prompt Tips and TemplatesTurn “Venice: Four Grok 4.6 Prompt Tips and Templates” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the source and model state before use.X: Matthew Berman's Creator Profile Card PromptTurn “X: Matthew Berman's Creator Profile Card Prompt” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the source and model state before use.X: Shawn's Grok 4.6 Single-Prompt 3D Ship CaseTurn “X: Shawn's Grok 4.6 Single-Prompt 3D Ship Case” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the source and model state before use.