The post cites Artificial Analysis's AA-Omniscience Non-Hallucination Rate:
Grok 4.5: 45.9%
Grok 4.6: 65.7%
GPT-5.6 Sol: 7.8%
GPT-5.6 Terra: 12.1%
GPT-5.6 Luna: 7.4%
The metric describes the proportion of cases in which a model acknowledges uncertainty rather than fabricating an answer when it does not know the answer. The post also cautions that Grok 4.6's accuracy still needs to be considered alongside this metric; a model should not be evaluated on refusal rate alone.
The author argues that Agent tasks compound decision errors over time. A model willing to say "I'm not sure" at critical points may be better suited to long, unsupervised tasks than one that always gives a confident answer. Some commenters shared more effective engineering practices: define acceptance criteria clearly, break work into smaller tasks, and have external tests or another model review the work step by step.
This is a community interpretation. The cited data and metric definition should be checked against Artificial Analysis's original methodology. Non-hallucination rate is not factual accuracy and may also be affected by refusal strategy.
Grok 4.6