This evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
xAI Official News · Read evidenceGrok 4.6 · Reviews and evidence
Which Grok 4.6 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Artificial Analysis · Read evidenceThis evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
BenchLM.ai · Read evidenceFull reviews and related reading
Selected evidence
Grok 4.6 Official Release: Benchmarks and Capability Evaluation
This evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- The xAI release table positions Grok 4.6 for long-horizon agents, coding, knowledge work, and interactive projects; competitor scores come from separate system cards or leaderboards with mismatched Terminal-Bench versions, reasoning tiers, and harnesses.
- Source boundary
- Supports the vendor’s stated positioning and conditional benchmark table when the software-engineering versus knowledge-work split is kept visible.
- Unsupported claims
- Does not support an independent retest, a unified ranking, current availability, or a production-performance guarantee.
Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6
This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- Artificial Analysis records Intelligence Index 61, Terminal-Bench v2.1 at 88.4%, and about $0.84 per task; full sample, reasoning tier, and private AA-Briefcase harness details are not public.
- Source boundary
- Supports a bounded intelligence/cost reading of this AA snapshot; it does not support merging v2.1 with xAI v3.0 or other harness scores.
- Unsupported claims
- Does not support current pricing, Tabbit availability, a general success rate, or production-safety guarantees.
BenchLM: Grok 4.6's Public Scores, Speed, and Cost
This evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- The BenchLM snapshot records 63.4/100, about 66 tokens/s, and roughly 32.3 seconds to first token; evidence is insufficient for several categories including Reasoning, Knowledge, and Math.
- Source boundary
- Supports reading speed, TTFT, and the composite only within the page’s evidence coverage, and helps define retest questions.
- Unsupported claims
- Does not support treating the composite as a complete capability profile or making claims about later prices or unmeasured categories.
Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task
This evidence note covers “Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- In Cursor, Grok 4.6 Extra High and GPT-5.6 Sol Medium ran the same backend plan over roughly 2,500 lines of code, with Fable 5 High as reviewer; it was one task.
- Source boundary
- Supports one same-start engineering case that motivates checking cost, context contamination, and complex-logic boundaries.
- Unsupported claims
- Does not support an overall ranking, a statistical success rate, or cross-Cursor/Codex conclusions; subscription usage is not API cost.
All sources
All sources
Grok 4.6 Official Release: Benchmarks and Capability Evaluation
This evidence note covers “Grok 4.6 Official Release: Benchmarks and Capability Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- The xAI release table positions Grok 4.6 for long-horizon agents, coding, knowledge work, and interactive projects; competitor scores come from separate system cards or leaderboards with mismatched Terminal-Bench versions, reasoning tiers, and harnesses.
- Source boundary
- Supports the vendor’s stated positioning and conditional benchmark table when the software-engineering versus knowledge-work split is kept visible.
- Unsupported claims
- Does not support an independent retest, a unified ranking, current availability, or a production-performance guarantee.
Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6
This evidence note covers “Artificial Analysis: Intelligence and Cost Evaluation of Grok 4.6” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- Artificial Analysis records Intelligence Index 61, Terminal-Bench v2.1 at 88.4%, and about $0.84 per task; full sample, reasoning tier, and private AA-Briefcase harness details are not public.
- Source boundary
- Supports a bounded intelligence/cost reading of this AA snapshot; it does not support merging v2.1 with xAI v3.0 or other harness scores.
- Unsupported claims
- Does not support current pricing, Tabbit availability, a general success rate, or production-safety guarantees.
BenchLM: Grok 4.6's Public Scores, Speed, and Cost
This evidence note covers “BenchLM: Grok 4.6's Public Scores, Speed, and Cost” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- The BenchLM snapshot records 63.4/100, about 66 tokens/s, and roughly 32.3 seconds to first token; evidence is insufficient for several categories including Reasoning, Knowledge, and Math.
- Source boundary
- Supports reading speed, TTFT, and the composite only within the page’s evidence coverage, and helps define retest questions.
- Unsupported claims
- Does not support treating the composite as a complete capability profile or making claims about later prices or unmeasured categories.
Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task
This evidence note covers “Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- In Cursor, Grok 4.6 Extra High and GPT-5.6 Sol Medium ran the same backend plan over roughly 2,500 lines of code, with Fable 5 High as reviewer; it was one task.
- Source boundary
- Supports one same-start engineering case that motivates checking cost, context contamination, and complex-logic boundaries.
- Unsupported claims
- Does not support an overall ranking, a statistical success rate, or cross-Cursor/Codex conclusions; subscription usage is not API cost.
Emergent: Breaking Down Grok 4.6's Evaluation Results
Emergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61, tied with GPT-5.6 Sol; but on DeepSWE 。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “Emergent: Breaking Down Grok 4.6's Evaluation Results”; do not merge other versions or reasoning tiers.
- Task and harness
- Emergent argues that Grok 4.6's public results show a clear capability distribution: it is strong on knowledge-work evaluations but relatively weaker on pure software-engineering evaluations. Its Intelligence Index is 61 The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
KIE: Grok 4.6 Release Evaluation and Capability Breakdown
Article conclusion KIE summarizes Grok 4.6's Intelligence Index as 61, tied with GPT-5.6 Sol; its price is $2 for input and $6 for output per 1M tokens. The article argues that its main selling point is its price-to-intelligence ratio, rather than leading on e。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “KIE: Grok 4.6 Release Evaluation and Capability Breakdown”; do not merge other versions or reasoning tiers.
- Task and harness
- Article conclusion KIE summarizes Grok 4.6's Intelligence Index as 61, tied with GPT-5.6 Sol; its price is $2 for input and $6 for output per 1M tokens. The article argues that its main selling point is its price-to-inte The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Medium Line-by-Line Analysis: Grok 4.6 Compared with Sol and Fable
Article conclusion Based on the ten-row evaluation table released by xAI, the author counted each result: among the 9 rows with scores for both Grok 4.6 and GPT-5.6 Sol Max, Grok 4.6 won 6 and lost only on DeepSWE v1.1 and Terminal-Bench v3.0; in the ten-row c。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “Medium Line-by-Line Analysis: Grok 4.6 Compared with Sol and Fable”; do not merge other versions or reasoning tiers.
- Task and harness
- Article conclusion Based on the ten-row evaluation table released by xAI, the author counted each result: among the 9 rows with scores for both Grok 4.6 and GPT-5.6 Sol Max, Grok 4.6 won 6 and lost only on DeepSWE v1.1 a The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Reddit r/cursor: Community Discussion of the Gap Between Grok 4.6 Leaderboards and User Experience
Key points from the discussion Some users believe that Grok 4.6's upgrade over 4.5 is consistent with leaderboard trends, especially on Agentic and coding tasks; others believe that Grok 4.5's past public scores did not match their actual experience, and there。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “Reddit r/cursor: Community Discussion of the Gap Between Grok 4.6 Leaderboards and User Experience”; do not merge other versions or reasoning tiers.
- Task and harness
- Key points from the discussion Some users believe that Grok 4.6's upgrade over 4.5 is consistent with leaderboard trends, especially on Agentic and coding tasks; others believe that Grok 4.5's past public scores did not The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Reddit r/cursor: Grok 4.6's Non-Hallucination Rate and Refusal Calibration
Data cited in the article The post cites Artificial Analysis's AA-Omniscience Non-Hallucination Rate: GPT-5.6 Terra: 12.1%。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “Reddit r/cursor: Grok 4.6's Non-Hallucination Rate and Refusal Calibration”; do not merge other versions or reasoning tiers.
- Task and harness
- Data cited in the article The post cites Artificial Analysis's AA-Omniscience Non-Hallucination Rate: GPT-5.6 Terra: 12.1%。 The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench
The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73%.。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench”; do not merge other versions or reasoning tiers.
- Task and harness
- The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73% The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Reddit r/singularity: Grok 4.6 Benchmarks and Real-World Coding Feedback
Key points from the post and comments The post primarily discusses real-world impressions in light of xAI/Artificial Analysis scorecards. A highly engaged comment describes Grok 4.6 as "cheap and fast" for coding and shares a workflow in which Opus handles pla。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “Reddit r/singularity: Grok 4.6 Benchmarks and Real-World Coding Feedback”; do not merge other versions or reasoning tiers.
- Task and harness
- Key points from the post and comments The post primarily discusses real-world impressions in light of xAI/Artificial Analysis scorecards. A highly engaged comment describes Grok 4.6 as "cheap and fast" for coding and sha The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
X / TypingMind: Comparing Grok 4.6 with the Same Paper-Cut Animation Prompt
TypingMind said it gave the same “traditional Chinese paper-cut-style animation” prompt to Grok 4.6, GPT-5.6 Sol, Claude Opus 5, and Qwen 3.8 Max to compare the generated results. The post included a video and, in follow-up replies, provided the complete promp。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “X / TypingMind: Comparing Grok 4.6 with the Same Paper-Cut Animation Prompt”; do not merge other versions or reasoning tiers.
- Task and harness
- TypingMind said it gave the same “traditional Chinese paper-cut-style animation” prompt to Grok 4.6, GPT-5.6 Sol, Claude Opus 5, and Qwen 3.8 Max to compare the generated results. The post included a video and, in follow The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
X: Matthew Berman's Same-Prompt Comparison of Profile Cards
The author gave Grok 4.6, GPT-5.6 Sol, and Fable 5 the same prompt, asking them to generate a social-app profile card for a creator and comparing the results.。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “X: Matthew Berman's Same-Prompt Comparison of Profile Cards”; do not merge other versions or reasoning tiers.
- Task and harness
- The author gave Grok 4.6, GPT-5.6 Sol, and Fable 5 the same prompt, asking them to generate a social-app profile card for a creator and comparing the results.。 The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
X: Mike P's Grok 4.6 vs. Opus 5 on a Long-Data Fitness Report Task
The author gave the model complete Garmin and Apple Health data from 2019 to the present and asked it to generate a fitness report covering all types of exercise. The report was to focus on the author's cycling history, track long-term metrics and progress, an。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “X: Mike P's Grok 4.6 vs. Opus 5 on a Long-Data Fitness Report Task”; do not merge other versions or reasoning tiers.
- Task and harness
- The author gave the model complete Garmin and Apple Health data from 2019 to the present and asked it to generate a fitness report covering all types of exercise. The report was to focus on the author's cycling history, The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
X: Same-Prompt Cost and Speed Comparison of Grok 4.6 and Opus 5
The author says they tested Grok 4.6 and Opus 5 with the same prompt: Grok 4.6: approximately $1.74, completed in about 5 minutes.。
Unverified: the original source could not be rechecked.
- Model and version
- Grok 4.6; source title “X: Same-Prompt Cost and Speed Comparison of Grok 4.6 and Opus 5”; do not merge other versions or reasoning tiers.
- Task and harness
- The author says they tested Grok 4.6 and Opus 5 with the same prompt: Grok 4.6: approximately $1.74, completed in about 5 minutes.。 The complete task set and runtime parameters are not fully public.
- Sample and date
- Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.
Grok 4.6
Compare Grok 4.6 in Tabbit
Model access, features, and permissions depend on your current client account.