Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Grok 4.6 · Community source · Personal experience

Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench

The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73%.。

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model and version
Grok 4.6; source title “Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench”; do not merge other versions or reasoning tiers.
Task and harness
The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73% The complete task set and runtime parameters are not fully public.
Sample and date
Source note reviewed 2026-09-20; sample count, repeats, and raw logs remain unknown where undisclosed.

Key data and applicable tasks

Scores discussed

The comments cite a 65.9% score for Grok 4.6 on DeepSWE v1.1, noting that it is higher than DeepSeek V4 Pro 0813's 62.7%. This indicates a clear improvement over Grok 4.5's 54%, but it remains below GPT-5.6 Sol Max's 73%.

The post also discusses version changes in Terminal-Bench: after the move from 2.1 to 3.0, models that had originally scored close to 90% could fall back to around 30%. Scores from different versions therefore cannot be compared directly.

User feedback

  • One user uses Grok 4.6 in Cursor Enterprise and considers it suitable for work involving large numbers of tokens.

  • One user considers Grok 4.6's capabilities close to Opus 5's, but says the 500K context window remains limiting for some architect-type work.

  • Another group of comments uses Grok 4.6 for simple E2E tasks and sub-Agents, describing it as "smart enough and very fast."

  • Some comments argue that DeepSWE reflects only repository tasks and hidden tests, and cannot represent architecture, maintainability, or long-term collaboration in real development.

Conclusion

This post is more of a quick community interpretation of public scores than an independent rerun. Its value lies in reminding readers to consider DeepSWE, the Terminal-Bench version, and the test objective together when assessing Grok 4.6's coding results, and to evaluate "speed suitable for sub-Agents" separately from "stability when completing complex engineering work independently."

What this supports

  • Supports reading the task observation or editorial conclusion in “Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench” under the stated source conditions.

What this does not support

  • Does not support extending “Reddit r/opencodeCLI: Discussion of Grok 4.6 on DeepSWE and Terminal-Bench” to a universal ranking or production guarantee; its task set, runtime parameters, and independent repeats are limited or undisclosed.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit r/opencodeCLI · u/minxio · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Grok 4.6

Compare Grok 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Grok 4.6: What Changed, What It Costs, and Who It Fits

A sourced guide to Grok 4.6: the 500K context model, benchmark-version split, live API price and a safer pilot decision.

Comparison · English

Grok 4.7 vs Grok 4.6: Same Rate, Longer Bills

Grok 4.7 lists the same $2/$6 API rate and 500K context as Grok 4.6. At xhigh it used about 81k output tokens per intelligence task, versus 36k.

Related reviews

Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same TaskThis evidence note covers “Reddit r/cursor: Grok 4.6 vs. GPT-5.6 Sol on the Same Task” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.X: Matthew Berman's Same-Prompt Comparison of Profile CardsThe author gave Grok 4.6, GPT-5.6 Sol, and Fable 5 the same prompt, asking them to generate a social-app profile card for a creator and comparing the results.。X: Mike P's Grok 4.6 vs. Opus 5 on a Long-Data Fitness Report TaskThe author gave the model complete Garmin and Apple Health data from 2019 to the present and asked it to generate a fitness report covering all types of exercise. The report was to focus on the author's cycling history, track long-term metrics and progress, an。X: Same-Prompt Cost and Speed Comparison of Grok 4.6 and Opus 5The author says they tested Grok 4.6 and Opus 5 with the same prompt: Grok 4.6: approximately $1.74, completed in about 5 minutes.。X: Matthew Berman's Creator Profile Card PromptTurn “X: Matthew Berman's Creator Profile Card Prompt” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the source and model state before use.X: Shawn's Grok 4.6 Single-Prompt 3D Ship CaseTurn “X: Shawn's Grok 4.6 Single-Prompt 3D Ship Case” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the source and model state before use.X: Tim Jayas's Three.js Steam-Engine Prototype PromptTurn “X: Tim Jayas's Three.js Steam-Engine Prototype Prompt” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the source and model state before use.Venice: Four Grok 4.6 Prompt Tips and TemplatesTurn “Venice: Four Grok 4.6 Prompt Tips and Templates” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the source and model state before use.