Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaQwen3.8 Max

Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination Cost

Original source

NYU Shanghai RITS

AuthorRITS editorial team (no individual author listed in the body; the end of the article notes staff review)

Tabbit curation2026-08-19

Read original

One-sentence takeaway

RITS's organization of the Artificial Analysis snapshot shows that Qwen3.8-Max is close to the top models on agentic capability, but gets there mainly by taking longer Agent trajectories: about 64 turns per GDPval-AA task at roughly $1.14, while AA-LCR and AA-Omniscience declined and the observed hallucination rate rose from 23% to 40%.

Use cases

  • Suitable tasks: Research, office, and coding Agents that can tolerate longer runtimes and need multi-turn tool calls and sustained evidence collection; evaluating whether completion quality justifies the extra turns in a cost model.

  • Unsuitable tasks: Low-latency customer service, fixed token budgets, and factual question answering with a hard hallucination-rate threshold.

  • Applicable model versions: Qwen3.8-Max GA and the version evaluated by Artificial Analysis; the article mixes 53/56/58-point snapshots from different dates, so they must not be treated as a constant for one version.

  • Applicable clients, Agents, or APIs: Artificial Analysis's independent evaluation path; the article does not provide a directly reusable Qwen API configuration.

  • Recommended reasoning tier and parameters: Not disclosed; reproduction requires fixing the endpoint, reasoning settings, maximum turns, and tool permissions.

Test environment and workflow steps

  1. RITS compiled public snapshots of the Artificial Analysis Intelligence Index v4.1 and distinguished the full index from the Agentic Index; the page says the index includes GDPval-AA, Terminal-Bench, τ³-Banking, HLE, SciCode, GPQA, CritPt, AA-LCR, and AA-Omniscience, among other evaluations.

  2. The main task is GDPval-AA: work tasks from 44 occupations, with shell and web tools allowed; the article gives changes in the model's Elo, turns per task, and cost.

  3. It compares Qwen3.7-Max, Qwen3.8-Max, Kimi K3, GPT-5.6 Sol Max, and Claude Opus 5; the article also warns that index versions and endpoint issues can change snapshots.

  4. A review should record index scores separately from task trajectories: score, turns, cost, hallucination rate, and evaluations that incurred deductions must not be compressed into a single conclusion that one model is “smarter.”

Original evidence and data

  • Index evolution: The article records Artificial Analysis's total score for Qwen3.8-Max moving from 53 (withdrawn because of intermittent problems with the evaluated endpoint) to 56, then revised to 58; publication must therefore state the snapshot date and index version.

  • GDPval-AA Elo: Qwen3.8-Max scored 1,739; Qwen3.7-Max scored 1,271, a difference of 468; GPT-5.6 Sol Max scored 1,730; Kimi K3 scored 1,685; Claude Opus 5 scored 1,852.

  • Trajectory length: Qwen3.8-Max used about 64 turns per task, versus about 14 turns for Qwen3.7-Max.

  • Cost: The article records about $1.14 per task for the Intelligence Index, versus about $0.53 for Qwen3.7-Max and about $0.86 for Kimi K3; the current Artificial Analysis model page shows about $1.13, which is the same order of magnitude but not the same collection snapshot.

  • Regression signals: AA-LCR fell by 2 points; AA-Omniscience fell by 10 points; the article records the observed hallucination rate rising from 23% to 40%.

  • Agentic Index position: The article describes Qwen3.8-Max's 58 points as comparable to Claude Opus 5 in its xhigh tier, and 1 point below Opus 5 max; this is a particular Artificial Analysis index, not a universal ranking across all tasks.

Conclusions

The most reusable conclusion is not “Qwen3.8-Max has surpassed Opus 5,” but rather “it is willing to keep working longer on agentic tasks, thereby approaching top-tier scores while taking on greater token, latency, and hallucination risks.” Use budget guardrails: set a maximum number of turns, a cost ceiling, and factual verification before letting the model expand its search; when reliability like AA-LCR/Omniscience matters more than completion rate, add a second model or human review.

Limitations

  • This is RITS's secondary organization of Artificial Analysis results, not a complete benchmark run by RITS; readers should return to the Artificial Analysis model page and methodology page for verification.

  • The article cites index snapshots from multiple dates; 53, 56, and 58 cannot be treated as values comparable at one point in time.

  • GDPval's turns, cost, and hallucination rate depend on the specific harness, endpoint, and index version; they cannot be directly extrapolated to Qwen Studio, OpenCode, or third-party providers.

  • The article does not disclose the complete input prompt, temperature, maximum tokens, tool schema, or raw output for each question, so only the conclusions and metric definitions can be reviewed; the experiment cannot be fully rerun from the article alone.

Reproduction steps

  1. Record the current Artificial Analysis model page, Agentic Index page, index version, and collection time.

  2. Rerun a small GDPval-style task set with a fixed Qwen3.8-Max endpoint, fixing the tool set, maximum turns, reasoning effort, temperature, and output limit.

  3. For each task, record success/failure, turns, tool calls, input/output/reasoning tokens, cost, factual errors, and human correction time.

  4. Set maximums of 14 and 64 turns separately and compare whether the pass-rate improvement from the extra turns is worth the cost; label the result as your own reproduction experiment and do not present it as Artificial Analysis's original score.

Source excerpt or observation (short quote for compliance only)

The article summarizes the key change as “buying capability with inference budget”; this phrase corresponds to its turn, cost, and hallucination-rate data and cannot be quoted separately from those numbers.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Qwen3.8 Max

Use and compare models in Tabbit

Qwen3.8 Max

Related reviews

MediaOfficial Qwen Blog2026-08-03

Qwen3.8-Max: Official Release Notes and Complete Performance Results

MediaArtificial Analysis

Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and Verbosity

MediaTrilogy AI Center of Excellence (Substack)2026-07-19

Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test

MediaBenchLM2026-08-17

Qwen3.8 Max: BenchLM's Source-Verifiable Benchmark Ledger

Qwen3.8 Max

Related prompts

Mediaqwen.ai2026-08-03

Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide

MediaEvoLink.AI Blog2026-08-03

Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide

CommunityReddit, r/QwenAI2026-08-09

Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission Boundaries

CommunityReddit, r/QwenAI2026-08-06

Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports