Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGrok 4.6

Grok 4.6 Official Release: Benchmarks and Capability Evaluation

Original source

xAI Official News

Source date2026-08-12

Tabbit curation2026-08-19

Read original

Executive summary

Grok 4.6 is positioned for long-running agent tasks, coding, knowledge work, and interactive/visual projects. The company says it ties with GPT-5.6 Sol on the Artificial Analysis Intelligence Index, with an overall score of 61; however, the official table also shows it performing better on knowledge-work evaluations while trailing GPT-5.6 Sol Max on software-engineering evaluations such as DeepSWE and Terminal-Bench.

Official evaluation table

EvaluationGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%—58.8%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Capabilities emphasized by the company

  • Maintain context over long-running tasks and continue making progress across research, repository operations, and application building.

  • Perform strongly on knowledge-work and legal-work evaluations such as GDPVal-AA v2, AA-Briefcase, and Harvey LAB.

  • Training covers general coding, knowledge work, kernel optimization, web development, CAD, and other agent environments.

  • The company says more self-testing and verification behavior appears in long trajectories.

Pricing and specifications

  • Context window: 500,000 tokens

  • Input price: $2 / 1M tokens

  • Output price: $6 / 1M tokens

  • Available through: Cursor, Grok Build, API, OpenRouter, Vercel, Cloudflare, and others

Limitations when reading the results

The company notes that competitor scores come from system cards or public leaderboards published by their respective developers, rather than from four models rerun in the same experimental environment. The table is therefore suitable for directional judgment, but should not be treated as a strict head-to-head experimental conclusion. In particular, the Terminal-Bench version, reasoning level, and execution harness all affect the results.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Grok 4.6

Use and compare models in Tabbit

Grok 4.6

Related reviews

MediaArtificial Analysis2026-08-12

Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6

MediaBenchLM.ai

BenchLM: Grok 4.6's Public Scores, Speed, and Cost

MediaEmergent Learn2026-08-13

Emergent: Breaking Down Grok 4.6's Evaluation Results

MediaKIE.ai Blog2026-08-13

KIE: Grok 4.6 Release Evaluation and Capability Breakdown

Grok 4.6

Related prompts

OfficialxAI Developers Docs

xAI Official Developer Documentation: Basic Prompts and Parameter Settings for Grok 4.6

MediaBuild Fast with AI2026-08-14

Build Fast with AI: General Methods from 100 Grok Prompts

MediaLayer3 Labs Resources

Layer3 Labs: Writing and Editing Prompt Methods for Grok 4.6

CommunityReddit r/LoveGrok

Reddit r/LoveGrok: Practical Project Instructions and Positive Constraints