Grok 4.6

Grok 4.6 · Reviews and evidence

Which Grok 4.6 conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

This evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

xAI Official News · Read evidence

This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Artificial Analysis · Read evidence

This evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

BenchLM.ai · Read evidence

Full reviews and related reading

Read the full analysis

Overview · English

Grok 4.6: What Changed, What It Costs, and Who It Fits

A sourced guide to Grok 4.6: the 500K context model, benchmark-version split, live API price and a safer pilot decision.

Selected evidence

OfficialVendor report

Grok 4.6 Official Release: Benchmarks and Capability Evaluation

This evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourcexAI Official News
Published2026-08-12
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
The xAI release table positions Grok 4.6 for long-horizon agents, coding, knowledge work, and interactive projects; competitor scores come from separate system cards or leaderboards with mismatched Terminal-Bench versions, reasoning tiers, and harnesses.
Source boundary
Supports the vendor’s stated positioning and conditional benchmark table when the software-engineering versus knowledge-work split is kept visible.
Unsupported claims
Does not support an independent retest, a unified ranking, current availability, or a production-performance guarantee.
AgentVisual generationInformation extractionReasoning
Media / benchmarkIndependent measurement

Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6

This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceArtificial Analysis
Published2026-08-12
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
Artificial Analysis records Intelligence Index 61, Terminal-Bench v2.1 at 88.4%, and about $0.84 per task; full sample, reasoning tier, and private AA-Briefcase harness details are not public.
Source boundary
Supports a bounded intelligence/cost reading of this AA snapshot; it does not support merging v2.1 with xAI v3.0 or other harness scores.
Unsupported claims
Does not support current pricing, Tabbit availability, a general success rate, or production-safety guarantees.
ReasoningCost
Media / benchmarkIndependent measurement

BenchLM: Grok 4.6's Public Scores, Speed, and Cost

This evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceBenchLM.ai
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
The BenchLM snapshot records 63.4/100, about 66 tokens/s, and roughly 32.3 seconds to first token; evidence is insufficient for several categories including Reasoning, Knowledge, and Math.
Source boundary
Supports reading speed, TTFT, and the composite only within the page’s evidence coverage, and helps define retest questions.
Unsupported claims
Does not support treating the composite as a complete capability profile or making claims about later prices or unmeasured categories.
AgentCostSpeed & latency
CommunityPersonal experience

Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task

This evidence note covers “Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit r/cursor
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
In Cursor, Grok 4.6 Extra High and GPT-5.6 Sol Medium ran the same backend plan over roughly 2,500 lines of code, with Fable 5 High as reviewer; it was one task.
Source boundary
Supports one same-start engineering case that motivates checking cost, context contamination, and complex-logic boundaries.
Unsupported claims
Does not support an overall ranking, a statistical success rate, or cross-Cursor/Codex conclusions; subscription usage is not API cost.
CodingAgent

All sources

All sources

15 / 15
OfficialVendor report

Grok 4.6 Official Release: Benchmarks and Capability Evaluation

This evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourcexAI Official News
Published2026-08-12
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
The xAI release table positions Grok 4.6 for long-horizon agents, coding, knowledge work, and interactive projects; competitor scores come from separate system cards or leaderboards with mismatched Terminal-Bench versions, reasoning tiers, and harnesses.
Source boundary
Supports the vendor’s stated positioning and conditional benchmark table when the software-engineering versus knowledge-work split is kept visible.
Unsupported claims
Does not support an independent retest, a unified ranking, current availability, or a production-performance guarantee.
AgentVisual generationInformation extractionReasoning
Media / benchmarkIndependent measurement

Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6

This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceArtificial Analysis
Published2026-08-12
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
Artificial Analysis records Intelligence Index 61, Terminal-Bench v2.1 at 88.4%, and about $0.84 per task; full sample, reasoning tier, and private AA-Briefcase harness details are not public.
Source boundary
Supports a bounded intelligence/cost reading of this AA snapshot; it does not support merging v2.1 with xAI v3.0 or other harness scores.
Unsupported claims
Does not support current pricing, Tabbit availability, a general success rate, or production-safety guarantees.
ReasoningCost
Media / benchmarkIndependent measurement

BenchLM: Grok 4.6's Public Scores, Speed, and Cost

This evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceBenchLM.ai
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
The BenchLM snapshot records 63.4/100, about 66 tokens/s, and roughly 32.3 seconds to first token; evidence is insufficient for several categories including Reasoning, Knowledge, and Math.
Source boundary
Supports reading speed, TTFT, and the composite only within the page’s evidence coverage, and helps define retest questions.
Unsupported claims
Does not support treating the composite as a complete capability profile or making claims about later prices or unmeasured categories.
AgentCostSpeed & latency
CommunityPersonal experience

Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task

This evidence note covers “Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit r/cursor
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
In Cursor, Grok 4.6 Extra High and GPT-5.6 Sol Medium ran the same backend plan over roughly 2,500 lines of code, with Fable 5 High as reviewer; it was one task.
Source boundary
Supports one same-start engineering case that motivates checking cost, context contamination, and complex-logic boundaries.
Unsupported claims
Does not support an overall ranking, a statistical success rate, or cross-Cursor/Codex conclusions; subscription usage is not API cost.
CodingAgent
Media / benchmarkIndependent measurement

Emergent: Breaking Down Grok 4.6's Evaluation Results

Emergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61, tied with GPT-5.6 Sol; but on DeepSWE 。

SourceEmergent Learn
Published2026-08-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “Emergent: Breaking Down Grok 4.6's Evaluation Results”; do not merge other versions or reasoning tiers.
Task and harness
Emergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61 The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
CodingInformation extractionReasoning
Media / benchmarkEditorial analysis

KIE: Grok 4.6 Release Evaluation and Capability Breakdown

Article conclusion KIE summarizes Grok 4.6's Intelligence Index as 61, tied with GPT-5.6 Sol; its price is $2 for input and $6 for output per 1M tokens. The article argues that its main selling point is its price-to-intelligence ratio, rather than leading on e。

SourceKIE.ai Blog
Published2026-08-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “KIE: Grok 4.6 Release Evaluation and Capability Breakdown”; do not merge other versions or reasoning tiers.
Task and harness
Article conclusion KIE summarizes Grok 4.6's Intelligence Index as 61, tied with GPT-5.6 Sol; its price is $2 for input and $6 for output per 1M tokens. The article argues that its main selling point is its price-to-inte The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
ReasoningCost
Media / benchmarkEditorial analysis

Medium Line-by-Line Analysis: Grok 4.6 Compared with Sol and Fable

Article conclusion Based on the ten-row evaluation table released by xAI, the author counted each result: among the 9 rows with scores for both Grok 4.6 and GPT-5.6 Sol Max, Grok 4.6 won 6 and lost only on DeepSWE v1.1 and Terminal-Bench v3.0; in the ten-row c。

SourceMedium / Data Science Collective
Published2026-08-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “Medium Line-by-Line Analysis: Grok 4.6 Compared with Sol and Fable”; do not merge other versions or reasoning tiers.
Task and harness
Article conclusion Based on the ten-row evaluation table released by xAI, the author counted each result: among the 9 rows with scores for both Grok 4.6 and GPT-5.6 Sol Max, Grok 4.6 won 6 and lost only on DeepSWE v1.1 a The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Reasoning
CommunityPersonal experience

Reddit r/cursor: Community Discussion of the Gap Between Grok 4.6 Leaderboards and User Experience

Key points from the discussion Some users believe that Grok 4.6's upgrade over 4.5 is consistent with leaderboard trends, especially on Agentic and coding tasks; others believe that Grok 4.5's past public scores did not match their actual experience, and there。

SourceReddit r/cursor
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “Reddit r/cursor: Community Discussion of the Gap Between Grok 4.6 Leaderboards and User Experience”; do not merge other versions or reasoning tiers.
Task and harness
Key points from the discussion Some users believe that Grok 4.6's upgrade over 4.5 is consistent with leaderboard trends, especially on Agentic and coding tasks; others believe that Grok 4.5's past public scores did not The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Agent
CommunityPersonal experience

Reddit r/cursor: Grok 4.6's Non-Hallucination Rate and Refusal Calibration

Data cited in the article The post cites Artificial Analysis's AA-Omniscience Non-Hallucination Rate: GPT-5.6 Terra: 12.1%。

SourceReddit r/cursor
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “Reddit r/cursor: Grok 4.6's Non-Hallucination Rate and Refusal Calibration”; do not merge other versions or reasoning tiers.
Task and harness
Data cited in the article The post cites Artificial Analysis's AA-Omniscience Non-Hallucination Rate: GPT-5.6 Terra: 12.1%。 The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Capability
CommunityPersonal experience

Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench

The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73%.。

SourceReddit r/opencodeCLI
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench”; do not merge other versions or reasoning tiers.
Task and harness
The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73% The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Coding
CommunityPersonal experience

Reddit r/singularity: Grok 4.6 Benchmarks and Real-World Coding Feedback

Key points from the post and comments The post primarily discusses real-world impressions in light of xAI/Artificial Analysis scorecards. A highly engaged comment describes Grok 4.6 as "cheap and fast" for coding and shares a workflow in which Opus handles pla。

SourceReddit r/singularity
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “Reddit r/singularity: Grok 4.6 Benchmarks and Real-World Coding Feedback”; do not merge other versions or reasoning tiers.
Task and harness
Key points from the post and comments The post primarily discusses real-world impressions in light of xAI/Artificial Analysis scorecards. A highly engaged comment describes Grok 4.6 as "cheap and fast" for coding and sha The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
AgentInformation extractionReasoningSpeed & latency
CommunityPersonal experience

X / TypingMind: Comparing Grok 4.6 with the Same Paper-Cut Animation Prompt

TypingMind said it gave the same “traditional Chinese paper-cut-style animation” prompt to Grok 4.6, GPT-5.6 Sol, Claude Opus 5, and Qwen 3.8 Max to compare the generated results. The post included a video and, in follow-up replies, provided the complete promp。

SourceX
Published2026-08-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “X / TypingMind: Comparing Grok 4.6 with the Same Paper-Cut Animation Prompt”; do not merge other versions or reasoning tiers.
Task and harness
TypingMind said it gave the same “traditional Chinese paper-cut-style animation” prompt to Grok 4.6, GPT-5.6 Sol, Claude Opus 5, and Qwen 3.8 Max to compare the generated results. The post included a video and, in follow The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Visual generationwriting
CommunityPersonal experience

X: Matthew Berman's Same-Prompt Comparison of Profile Cards

The author gave Grok 4.6, GPT-5.6 Sol, and Fable 5 the same prompt, asking them to generate a social-app profile card for a creator and comparing the results.。

SourceX
Published2026-08-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “X: Matthew Berman's Same-Prompt Comparison of Profile Cards”; do not merge other versions or reasoning tiers.
Task and harness
The author gave Grok 4.6, GPT-5.6 Sol, and Fable 5 the same prompt, asking them to generate a social-app profile card for a creator and comparing the results.。 The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
CodingVisual generationwriting
CommunityPersonal experience

X: Mike P's Grok 4.6 vs. Opus 5 on a Long-Data Fitness Report Task

The author gave the model complete Garmin and Apple Health data from 2019 to the present and asked it to generate a fitness report covering all types of exercise. The report was to focus on the author's cycling history, track long-term metrics and progress, an。

SourceX article
Published2026-08-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “X: Mike P's Grok 4.6 vs. Opus 5 on a Long-Data Fitness Report Task”; do not merge other versions or reasoning tiers.
Task and harness
The author gave the model complete Garmin and Apple Health data from 2019 to the present and asked it to generate a fitness report covering all types of exercise. The report was to focus on the author's cycling history, The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
CodingAgent
CommunityPersonal experience

X: Same-Prompt Cost and Speed Comparison of Grok 4.6 and Opus 5

The author says they tested Grok 4.6 and Opus 5 with the same prompt: Grok 4.6: approximately $1.74, completed in about 5 minutes.。

SourceX
Published2026-08-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
Grok 4.6; source title “X: Same-Prompt Cost and Speed Comparison of Grok 4.6 and Opus 5”; do not merge other versions or reasoning tiers.
Task and harness
The author says they tested Grok 4.6 and Opus 5 with the same prompt: Grok 4.6: approximately $1.74, completed in about 5 minutes.。 The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
CodingwritingCostSpeed & latency

Grok 4.6

Compare Grok 4.6 in Tabbit

Model access, features, and permissions depend on your current client account.