RITS's organization of the Artificial Analysis snapshot shows that Qwen3.8-Max is close to the top models on agentic capability, but gets there mainly by taking longer Agent trajectories: about 64 turns per GDPval-AA task at roughly $1.14, while AA-LCR and AA-Omniscience declined and the observed hallucination rate rose from 23% to 40%.
Suitable tasks: Research, office, and coding Agents that can tolerate longer runtimes and need multi-turn tool calls and sustained evidence collection; evaluating whether completion quality justifies the extra turns in a cost model.
Unsuitable tasks: Low-latency customer service, fixed token budgets, and factual question answering with a hard hallucination-rate threshold.
Applicable model versions: Qwen3.8-Max GA and the version evaluated by Artificial Analysis; the article mixes 53/56/58-point snapshots from different dates, so they must not be treated as a constant for one version.
Applicable clients, Agents, or APIs: Artificial Analysis's independent evaluation path; the article does not provide a directly reusable Qwen API configuration.
Recommended reasoning tier and parameters: Not disclosed; reproduction requires fixing the endpoint, reasoning settings, maximum turns, and tool permissions.
RITS compiled public snapshots of the Artificial Analysis Intelligence Index v4.1 and distinguished the full index from the Agentic Index; the page says the index includes GDPval-AA, Terminal-Bench, τ³-Banking, HLE, SciCode, GPQA, CritPt, AA-LCR, and AA-Omniscience, among other evaluations.
The main task is GDPval-AA: work tasks from 44 occupations, with shell and web tools allowed; the article gives changes in the model's Elo, turns per task, and cost.
It compares Qwen3.7-Max, Qwen3.8-Max, Kimi K3, GPT-5.6 Sol Max, and Claude Opus 5; the article also warns that index versions and endpoint issues can change snapshots.
A review should record index scores separately from task trajectories: score, turns, cost, hallucination rate, and evaluations that incurred deductions must not be compressed into a single conclusion that one model is “smarter.”
Index evolution: The article records Artificial Analysis's total score for Qwen3.8-Max moving from 53 (withdrawn because of intermittent problems with the evaluated endpoint) to 56, then revised to 58; publication must therefore state the snapshot date and index version.
GDPval-AA Elo: Qwen3.8-Max scored 1,739; Qwen3.7-Max scored 1,271, a difference of 468; GPT-5.6 Sol Max scored 1,730; Kimi K3 scored 1,685; Claude Opus 5 scored 1,852.
Trajectory length: Qwen3.8-Max used about 64 turns per task, versus about 14 turns for Qwen3.7-Max.
Cost: The article records about $1.14 per task for the Intelligence Index, versus about $0.53 for Qwen3.7-Max and about $0.86 for Kimi K3; the current Artificial Analysis model page shows about $1.13, which is the same order of magnitude but not the same collection snapshot.
Regression signals: AA-LCR fell by 2 points; AA-Omniscience fell by 10 points; the article records the observed hallucination rate rising from 23% to 40%.
Agentic Index position: The article describes Qwen3.8-Max's 58 points as comparable to Claude Opus 5 in its xhigh tier, and 1 point below Opus 5 max; this is a particular Artificial Analysis index, not a universal ranking across all tasks.
The most reusable conclusion is not “Qwen3.8-Max has surpassed Opus 5,” but rather “it is willing to keep working longer on agentic tasks, thereby approaching top-tier scores while taking on greater token, latency, and hallucination risks.” Use budget guardrails: set a maximum number of turns, a cost ceiling, and factual verification before letting the model expand its search; when reliability like AA-LCR/Omniscience matters more than completion rate, add a second model or human review.
This is RITS's secondary organization of Artificial Analysis results, not a complete benchmark run by RITS; readers should return to the Artificial Analysis model page and methodology page for verification.
The article cites index snapshots from multiple dates; 53, 56, and 58 cannot be treated as values comparable at one point in time.
GDPval's turns, cost, and hallucination rate depend on the specific harness, endpoint, and index version; they cannot be directly extrapolated to Qwen Studio, OpenCode, or third-party providers.
The article does not disclose the complete input prompt, temperature, maximum tokens, tool schema, or raw output for each question, so only the conclusions and metric definitions can be reviewed; the experiment cannot be fully rerun from the article alone.
Record the current Artificial Analysis model page, Agentic Index page, index version, and collection time.
Rerun a small GDPval-style task set with a fixed Qwen3.8-Max endpoint, fixing the tool set, maximum turns, reasoning effort, temperature, and output limit.
For each task, record success/failure, turns, tool calls, input/output/reasoning tokens, cost, factual errors, and human correction time.
Set maximums of 14 and 64 turns separately and compare whether the pass-rate improvement from the extra turns is worth the cost; label the result as your own reproduction experiment and do not present it as Artificial Analysis's original score.
The article summarizes the key change as “buying capability with inference budget”; this phrase corresponds to its turn, cost, and hallucination-rate data and cannot be quoted separately from those numbers.
Qwen3.8 Max