Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Grok 4.6 · Official source · Vendor report

Grok 4.6 Official Release: Benchmarks and Capability Evaluation

This evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Test conditions
The xAI release table positions Grok 4.6 for long-horizon agents, coding, knowledge work, and interactive projects; competitor scores come from separate system cards or leaderboards with mismatched Terminal-Bench versions, reasoning tiers, and harnesses.
Source boundary
Supports the vendor’s stated positioning and conditional benchmark table when the software-engineering versus knowledge-work split is kept visible.
Unsupported claims
Does not support an independent retest, a unified ranking, current availability, or a production-performance guarantee.

Key data and applicable tasks

Executive summary

Grok 4.6 is positioned for long-running agent tasks, coding, knowledge work, and interactive/visual projects. The company says it ties with GPT-5.6 Sol on the Artificial Analysis Intelligence Index, with an overall score of 61; however, the official table also shows it performing better on knowledge-work evaluations while trailing GPT-5.6 Sol Max on software-engineering evaluations such as DeepSWE and Terminal-Bench.

Official evaluation table

EvaluationGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%—58.8%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Capabilities emphasized by the company

  • Maintain context over long-running tasks and continue making progress across research, repository operations, and application building.

  • Perform strongly on knowledge-work and legal-work evaluations such as GDPVal-AA v2, AA-Briefcase, and Harvey LAB.

  • Training covers general coding, knowledge work, kernel optimization, web development, CAD, and other agent environments.

  • The company says more self-testing and verification behavior appears in long trajectories.

Pricing and specifications

  • Context window: 500,000 tokens

  • Input price: $2 / 1M tokens

  • Output price: $6 / 1M tokens

  • Available through: Cursor, Grok Build, API, OpenRouter, Vercel, Cloudflare, and others

Limitations when reading the results

The company notes that competitor scores come from system cards or public leaderboards published by their respective developers, rather than from four models rerun in the same experimental environment. The table is therefore suitable for directional judgment, but should not be treated as a strict head-to-head experimental conclusion. In particular, the Terminal-Bench version, reasoning level, and execution harness all affect the results.

What this supports

  • Supports the vendor’s stated positioning and conditional benchmark table when the software-engineering versus knowledge-work split is kept visible.

What this does not support

  • Does not support an independent retest, a unified ranking, current availability, or a production-performance guarantee.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

xAI Official News · Author not disclosed · Original publication date 2026-08-12 · Site edit date 2026-09-20

Open original source

Grok 4.6

Compare Grok 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Grok 4.6: What Changed, What It Costs, and Who It Fits

A sourced guide to Grok 4.6: the 500K context model, benchmark-version split, live API price and a safer pilot decision.

Comparison · English

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

Related reviews

Reddit r/singularity: Grok 4.6 Benchmarks and Real-World Coding FeedbackKey points from the post and comments The post primarily discusses real-world impressions in light of xAI/Artificial Analysis scorecards. A highly engaged comment describes Grok 4.6 as "cheap and fast" for coding and shares a workflow in which Opus handles pla。Emergent: Breaking Down Grok 4.6's Evaluation ResultsEmergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61, tied with GPT-5.6 Sol; but on DeepSWE 。Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.BenchLM: Grok 4.6's Public Scores, Speed, and CostThis evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.X: Nikolai Yakovenko on the Single-Prompt Playable-Game PatternThe author believes that making a playable game with a “one-sentence prompt” has become a common benchmark when a new model or Agent harness is released. The comments showed a case in which Grok 4.6 generated a simulation of Phantasy Star 1 and an AI Agent ope。X: Vaibhav Sisinty's Four-Task Prompt Frameworks for Grok 4.6The author lists four tasks worth trying immediately with Grok 4.6: Give the model a product idea and have it research the domain, build the app, design the UI, and deliver a working version from a single prompt.。X: Anshu's Grok 4.6 App and Design Workflow with the Same PromptThe author ran the same one-shot prompt in Grok Build with Grok 4.5 and Grok 4.6, then compared the resulting apps and designs. The author considers 4.6 a clear improvement over 4.5, saying it can even alternate for the lead with Fable on design tasks.。X: Matthew Berman's Creator Profile Card PromptTurn “X: Matthew Berman's Creator Profile Card Prompt” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the source and model state before use.