Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.8 Max · Media / benchmark · Independent measurement

Qwen3.8 Max: Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination Cost

NYU Shanghai RITS material discusses Qwen3.8 Max agent turns and hallucination/cost proxies; task set, tools, repeats, and version follow the disclosed portion and cannot generalize to every agent workload.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Study
NYU Shanghai RITS agentic index; version and harness follow the source
Metrics
Agent turns, hallucination, or cost proxies, not one quality score
Sample
Task set, repeats, and tools follow the disclosed portion
Boundary
Cannot generalize to every agent workload

Key data and applicable tasks

One-sentence takeaway

RITS's organization of the Artificial Analysis snapshot shows that Qwen3.8-Max is close to the top models on agentic capability, but gets there mainly by taking longer Agent trajectories: about 64 turns per GDPval-AA task at roughly $1.14, while AA-LCR and AA-Omniscience declined and the observed hallucination rate rose from 23% to 40%.

Use cases

  • Suitable tasks: Research, office, and coding Agents that can tolerate longer runtimes and need multi-turn tool calls and sustained evidence collection; evaluating whether completion quality justifies the extra turns in a cost model.

  • Unsuitable tasks: Low-latency customer service, fixed token budgets, and factual question answering with a hard hallucination-rate threshold.

  • Applicable model versions: Qwen3.8-Max GA and the version evaluated by Artificial Analysis; the article mixes 53/56/58-point snapshots from different dates, so they must not be treated as a constant for one version.

  • Applicable clients, Agents, or APIs: Artificial Analysis's independent evaluation path; the article does not provide a directly reusable Qwen API configuration.

  • Recommended reasoning tier and parameters: Not disclosed; reproduction requires fixing the endpoint, reasoning settings, maximum turns, and tool permissions.

Test environment and workflow steps

  1. RITS compiled public snapshots of the Artificial Analysis Intelligence Index v4.1 and distinguished the full index from the Agentic Index; the page says the index includes GDPval-AA, Terminal-Bench, τ³-Banking, HLE, SciCode, GPQA, CritPt, AA-LCR, and AA-Omniscience, among other evaluations.

  2. The main task is GDPval-AA: work tasks from 44 occupations, with shell and web tools allowed; the article gives changes in the model's Elo, turns per task, and cost.

  3. It compares Qwen3.7-Max, Qwen3.8-Max, Kimi K3, GPT-5.6 Sol Max, and Claude Opus 5; the article also warns that index versions and endpoint issues can change snapshots.

  4. A review should record index scores separately from task trajectories: score, turns, cost, hallucination rate, and evaluations that incurred deductions must not be compressed into a single conclusion that one model is “smarter.”

Original evidence and data

  • Index evolution: The article records Artificial Analysis's total score for Qwen3.8-Max moving from 53 (withdrawn because of intermittent problems with the evaluated endpoint) to 56, then revised to 58; publication must therefore state the snapshot date and index version.

  • GDPval-AA Elo: Qwen3.8-Max scored 1,739; Qwen3.7-Max scored 1,271, a difference of 468; GPT-5.6 Sol Max scored 1,730; Kimi K3 scored 1,685; Claude Opus 5 scored 1,852.

  • Trajectory length: Qwen3.8-Max used about 64 turns per task, versus about 14 turns for Qwen3.7-Max.

  • Cost: The article records about $1.14 per task for the Intelligence Index, versus about $0.53 for Qwen3.7-Max and about $0.86 for Kimi K3; the current Artificial Analysis model page shows about $1.13, which is the same order of magnitude but not the same collection snapshot.

  • Regression signals: AA-LCR fell by 2 points; AA-Omniscience fell by 10 points; the article records the observed hallucination rate rising from 23% to 40%.

  • Agentic Index position: The article describes Qwen3.8-Max's 58 points as comparable to Claude Opus 5 in its xhigh tier, and 1 point below Opus 5 max; this is a particular Artificial Analysis index, not a universal ranking across all tasks.

Conclusions

The most reusable conclusion is not “Qwen3.8-Max has surpassed Opus 5,” but rather “it is willing to keep working longer on agentic tasks, thereby approaching top-tier scores while taking on greater token, latency, and hallucination risks.” Use budget guardrails: set a maximum number of turns, a cost ceiling, and factual verification before letting the model expand its search; when reliability like AA-LCR/Omniscience matters more than completion rate, add a second model or human review.

Limitations

  • This is RITS's secondary organization of Artificial Analysis results, not a complete benchmark run by RITS; readers should return to the Artificial Analysis model page and methodology page for verification.

  • The article cites index snapshots from multiple dates; 53, 56, and 58 cannot be treated as values comparable at one point in time.

  • GDPval's turns, cost, and hallucination rate depend on the specific harness, endpoint, and index version; they cannot be directly extrapolated to Qwen Studio, OpenCode, or third-party providers.

  • The article does not disclose the complete input prompt, temperature, maximum tokens, tool schema, or raw output for each question, so only the conclusions and metric definitions can be reviewed; the experiment cannot be fully rerun from the article alone.

Reproduction steps

  1. Record the current Artificial Analysis model page, Agentic Index page, index version, and collection time.

  2. Rerun a small GDPval-style task set with a fixed Qwen3.8-Max endpoint, fixing the tool set, maximum turns, reasoning effort, temperature, and output limit.

  3. For each task, record success/failure, turns, tool calls, input/output/reasoning tokens, cost, factual errors, and human correction time.

  4. Set maximums of 14 and 64 turns separately and compare whether the pass-rate improvement from the extra turns is worth the cost; label the result as your own reproduction experiment and do not present it as Artificial Analysis's original score.

Source excerpt or observation (short quote for compliance only)

The article summarizes the key change as “buying capability with inference budget”; this phrase corresponds to its turn, cost, and hallucination-rate data and cannot be quoted separately from those numbers.

What this supports

  • Supports discussing trade-offs among agent turns, hallucination, and cost proxies.

What this does not support

  • Does not generalize across all toolchains, task sets, or production reliability.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

NYU Shanghai RITS · RITS editorial team (no individual author listed in the body; the end of the article notes staff review) · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Qwen3.8 Max

Compare Qwen3.8 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.8 Max: What Changed, What It Costs, and Who It Fits

A sourced Qwen3.8 Max overview covering the 0902 snapshot, multimodal boundary, benchmark caveats, access routes and a safer pilot.

Related reviews

Qwen3.8 Max: Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and VerbosityArtificial Analysis separates Qwen3.8 Max quality, cost, speed, and verbosity; page version, reasoning tier, provider, and task sample need a fresh check, and the aggregate index must not become a cross-version trend.Qwen3.8 Max: Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research BenchVals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elap。Qwen3.8 Max: Qwen3.8 Max: BenchLM's Source-Verifiable Benchmark LedgerBenchLM separates Qwen3.8 Max exact-source benchmark rows from its aggregate ranking; weights, providers, harnesses, samples, and dates differ, making it a verifiable ledger rather than a unified independent rerun.Qwen3.8 Max: Reddit Community: Qwen3.8-Max Coding Ability, Speed, and Usage QuotaThis is a community discussion asking whether Qwen3.8-Max is really suitable for programming. The feedback is polarized: some users consider it close to Claude/GPT, while others find it slow, expensive, and prone to overthinking. Another user used it to genera。Qwen3.8 Max: Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting GuideTurn Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling GuideTurn Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission BoundariesTurn Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission Boundaries into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” ReportsTurn Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.