Grok 4.6

Grok 4.6 review navigator

Official benchmarks, independent analysis, and community reports about Grok 4.6, clearly separated from Tabbit's own testing.

15 source-checked resourcesOfficial · Media · Community

Official

1 source-checked resources

Media

5 source-checked resources
MediaArtificial Analysis

Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6

Results at a glance Artificial Analysis rates Grok 4.6 at 61 on its Intelligence Index, tied with GPT-5.6 Sol and behind only Claude Opus 5 and Fable 5. The firm's interpretation is that Grok 4.6's advantage is not limited to static reasoning; it lies in its o。

MediaBenchLM.ai

BenchLM: Grok 4.6's Public Scores, Speed, and Cost

BenchLM overall public score: 63.4 / 100. Public leaderboard rank: 43 / 218. Agentic category: 7 / 130. Coding category: 18 / 135.。

MediaEmergent Learn

Emergent: Breaking Down Grok 4.6's Evaluation Results

Emergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61, tied with GPT-5.6 Sol; but on DeepSWE 。

MediaKIE.ai Blog

KIE: Grok 4.6 Release Evaluation and Capability Breakdown

Article conclusion KIE summarizes Grok 4.6's Intelligence Index as 61, tied with GPT-5.6 Sol; its price is $2 for input and $6 for output per 1M tokens. The article argues that its main selling point is its price-to-intelligence ratio, rather than leading on e。

MediaMedium / Data Science Collective

Medium Line-by-Line Analysis: Grok 4.6 Compared with Sol and Fable

Article conclusion Based on the ten-row evaluation table released by xAI, the author counted each result: among the 9 rows with scores for both Grok 4.6 and GPT-5.6 Sol Max, Grok 4.6 won 6 and lost only on DeepSWE v1.1 and Terminal-Bench v3.0; in the ten-row c。

Community

9 source-checked resources
CommunityReddit r/cursor

Reddit r/cursor: Community Discussion of the Gap Between Grok 4.6 Leaderboards and User Experience

Key points from the discussion Some users believe that Grok 4.6's upgrade over 4.5 is consistent with leaderboard trends, especially on Agentic and coding tasks; others believe that Grok 4.5's past public scores did not match their actual experience, and there。

CommunityReddit r/cursor

Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task

In Cursor, the author used Grok 4.6 Extra High and GPT-5.6 Sol Medium to execute the same detailed backend plan, starting from the same point and using the same plan, with a task size of approximately 2,500 lines of code; Fable 5 High served as an independent 。

CommunityReddit r/cursor

Reddit r/cursor: Grok 4.6's Non-Hallucination Rate and Refusal Calibration

Data cited in the article The post cites Artificial Analysis's AA-Omniscience Non-Hallucination Rate: GPT-5.6 Terra: 12.1%。

CommunityReddit r/opencodeCLI

Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench

The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73%.。

CommunityReddit r/singularity

Reddit r/singularity: Grok 4.6 Benchmarks and Real-World Coding Feedback

Key points from the post and comments The post primarily discusses real-world impressions in light of xAI/Artificial Analysis scorecards. A highly engaged comment describes Grok 4.6 as "cheap and fast" for coding and shares a workflow in which Opus handles pla。

CommunityX

X / TypingMind: Comparing Grok 4.6 with the Same Paper-Cut Animation Prompt

TypingMind said it gave the same “traditional Chinese paper-cut-style animation” prompt to Grok 4.6, GPT-5.6 Sol, Claude Opus 5, and Qwen 3.8 Max to compare the generated results. The post included a video and, in follow-up replies, provided the complete promp。

CommunityX

X: Matthew Berman's Same-Prompt Comparison of Profile Cards

The author gave Grok 4.6, GPT-5.6 Sol, and Fable 5 the same prompt, asking them to generate a social-app profile card for a creator and comparing the results.。

CommunityX article

X: Mike P's Grok 4.6 vs. Opus 5 on a Long-Data Fitness Report Task

The author gave the model complete Garmin and Apple Health data from 2019 to the present and asked it to generate a fitness report covering all types of exercise. The report was to focus on the author's cycling history, track long-term metrics and progress, an。

CommunityX

X: Same-Prompt Cost and Speed Comparison of Grok 4.6 and Opus 5

The author says they tested Grok 4.6 and Opus 5 with the same prompt: Grok 4.6: approximately $1.74, completed in about 5 minutes.。

Grok 4.6

Use and compare models in Tabbit

Official benchmarks, independent analysis, and community reports about Grok 4.6, clearly separated from Tabbit's own testing.