Grok 4.6 · Community source · Personal experience
Data cited in the article The post cites Artificial Analysis's AA-Omniscience Non-Hallucination Rate: GPT-5.6 Terra: 12.1%。
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
The post cites Artificial Analysis's AA-Omniscience Non-Hallucination Rate:
Grok 4.5: 45.9%
Grok 4.6: 65.7%
GPT-5.6 Sol: 7.8%
GPT-5.6 Terra: 12.1%
GPT-5.6 Luna: 7.4%
The metric describes the proportion of cases in which a model acknowledges uncertainty rather than fabricating an answer when it does not know the answer. The post also cautions that Grok 4.6's accuracy still needs to be considered alongside this metric; a model should not be evaluated on refusal rate alone.
The author argues that Agent tasks compound decision errors over time. A model willing to say "I'm not sure" at critical points may be better suited to long, unsupervised tasks than one that always gives a confident answer. Some commenters shared more effective engineering practices: define acceptance criteria clearly, break work into smaller tasks, and have external tests or another model review the work step by step.
This is a community interpretation. The cited data and metric definition should be checked against Artificial Analysis's original methodology. Non-hallucination rate is not factual accuracy and may also be affected by refusal strategy.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Reddit r/cursor · u/KaiThoughtArchitect · Original publication date Unknown · Site edit date 2026-09-20
Open original sourceGrok 4.6
Download the Tabbit client to check model access