Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.8 Max · Media / benchmark · Independent measurement

Qwen3.8 Max: Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and Verbosity

Artificial Analysis separates Qwen3.8 Max quality, cost, speed, and verbosity; page version, reasoning tier, provider, and task sample need a fresh check, and the aggregate index must not become a cross-version trend.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Platform/version
Artificial Analysis index page; version and window need a fresh check
Metrics
Quality, cost, speed, and verbosity are interpreted separately
Configuration
Provider, tier, sample, and task set are incomplete
Boundary
Do not combine the aggregate index into a cross-version trend

Key data and applicable tasks

One-sentence takeaway

Artificial Analysis's current page gives Qwen3.8 Max an Intelligence Index score of 58 (No. 10 among 180 comparable models), while also recording its low speed of about 45 token/s, roughly 150M total output tokens across the index evaluations, and an approximate per-task cost of $1.13. It therefore looks more like a “high-quality but slow and highly deliberative” Agent model than a low-latency chat model.

Use cases

  • Suitable tasks: Agent tasks requiring tool calls, long-chain delivery, code execution, and high completion quality; model selection that compares models by task cost rather than unit price alone.

  • Unsuitable tasks: Highly real-time conversations, high-throughput services with short answers, and scenarios that treat a text-based English composite index as a direct proxy for multimodal or Chinese-language business quality.

  • Applicable model version: The reasoning version of Qwen3.8 Max; page data may change with the provider, evaluation version, and date.

  • Applicable client, Agent, or API: Artificial Analysis's unified evaluation environment; this does not equal actual latency in Qwen Studio, DashScope, or third-party routers.

  • Recommended reasoning tier and parameters: The complete prompt for a reproducible single-model test is not public; the methodology page generally uses temperature 0.6 for reasoning models and the maximum output limit provided by the model. Before launch, test low, medium, and high reasoning tiers separately in your own Agent harness.

Test environment and workflow

  1. Artificial Analysis Intelligence Index v4.1.1 includes GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, Humanity’s Last Exam, GPQA Diamond, and CritPt, for nine evaluations in total.

  2. The index is weighted as Agents 34%, Coding 24%, Scientific Reasoning 24%, and General 18%; different evaluations use different numbers of repetitions and tool configurations.

  3. The public methodology page gives sample sizes including: GDPval-AA v2, 220 tasks, 1 run; τ³-Banking, 97 questions, 5 runs; Terminal-Bench v2.1, 89 tasks, 3 runs; SciCode, 288 subquestions, 3 runs; AA-LCR, 100 questions, 3 runs; AA-Omniscience, 6,000 questions, 1 run; HLE, 2,158 questions, 1 run; GPQA Diamond, 198 questions, 5 runs; and CritPt, 70 questions, 5 runs.

  4. The methodology page states that the evaluations use zero-shot instructions, unified scoring, and pass@1; GDPval-AA v2 uses the Stirrup harness, an E2B sandbox, file output, and web/code tools.

  5. For review, open both the model page and the Artificial Analysis Intelligence Index methodology page, and record the page date, evaluation version, price, and model endpoint; do not screenshot only the composite score.

Raw evidence and data

  • Overall quality: 58; the model summary on the page shows No. 10 among 180 comparable models.

  • Cost: $2.00 / 1M tokens input, $6.00 / 1M tokens output, 88% cache discount; the Artificial Analysis page shows an approximate Intelligence Index per-task cost of $1.13 and a full-index evaluation cost of approximately $1,741.41.

  • Speed: Output speed of about 44.6 token/s, listed by the page as No. 117/180 among comparable models; the summary places it in speed tier 1/4.

  • Output volume: Approximately 150M tokens generated cumulatively across the index evaluations; the page places it at No. 70/180 among comparable models for verbosity and explicitly warns that the model is very verbose.

  • Input capabilities: Supports text, image, and video input, with text output; the context window is listed as 1M tokens.

  • Methodological boundary: The methodology page describes the Intelligence Index as an English, text-only evaluation; visual, audio, and multilingual capabilities are measured separately and are not included in the composite score.

Conclusion

This source supports three conclusions that hold simultaneously: Qwen3.8-Max's overall capability is near the top; its task cost is not determined only by the $2/$6 unit prices, but is amplified by reasoning and output volume; and in human use, waiting time and answer length may become bottlenecks before the overall score does. When deploying it as the primary model for high-value Agents, treat “successful cost per task, completion latency, number of tool calls, and output tokens” as one acceptance sheet rather than looking only at the Intelligence Index.

Limitations

  • The page is a dynamic leaderboard; rankings, index versions, providers, and prices may change, so data collected on one day does not represent a permanent ranking.

  • The Intelligence Index is a weighted aggregate score focused on English text and cannot directly represent image, video, Chinese-language, or domain-specific tasks.

  • Artificial Analysis's methodology page publishes the evaluation structure and parameters, but the complete private inputs, execution traces, and details of every model endpoint are not all shown on the model page; readers cannot fully reproduce a score of 58 from the page alone.

  • Per-task cost includes weighted input, cache, reasoning, and answer tokens; it should not be treated as the quote for one ordinary API request.

Reproduction steps

  1. Fix the collection date and save the model page's score, ranking, input/output prices, speed, context, and output-volume fields.

  2. Fix Artificial Analysis Intelligence Index v4.1.1 and the weights of its nine evaluations; do not mix data from later versions into the same table.

  3. Build a zero-shot regression set of the same tasks on your own Qwen3.8-Max endpoint, recording at least success rate, time to first token, total elapsed time, reasoning tokens, answer tokens, tool-call count, and cost per task.

  4. Rerun with low/medium/high reasoning effort, compare “quality improvement ÷ additional tokens” and “successful-task cost,” and do not use this to claim that you reproduced Artificial Analysis's absolute score.

Source excerpt or observation (short quotation for compliance only)

The page summary describes Qwen3.8 Max as “notably slow and very verbose”; the methodology page also states that the Intelligence Index is only a composite comparison metric and cannot be applied directly to every use case. The two points should be cited together.

What this supports

  • Supports separate comparison of public quality, cost, speed, and verbosity dimensions.

What this does not support

  • Does not make the aggregate index a fixed-version production quality or trend.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis evaluation team · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Qwen3.8 Max

Compare Qwen3.8 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.8 Max: What Changed, What It Costs, and Who It Fits

A sourced Qwen3.8 Max overview covering the 0902 snapshot, multimodal boundary, benchmark caveats, access routes and a safer pilot.

Related reviews

Qwen3.8 Max: Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination CostNYU Shanghai RITS material discusses Qwen3.8 Max agent turns and hallucination/cost proxies; task set, tools, repeats, and version follow the disclosed portion and cannot generalize to every agent workload.Qwen3.8 Max: Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research BenchVals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elap。Qwen3.8 Max: Reddit Community: Qwen3.8-Max Coding Ability, Speed, and Usage QuotaThis is a community discussion asking whether Qwen3.8-Max is really suitable for programming. The feedback is polarized: some users consider it close to Claude/GPT, while others find it slow, expensive, and prone to overthinking. Another user used it to genera。Qwen3.8 Max: Qwen3.8 Max: BenchLM's Source-Verifiable Benchmark LedgerBenchLM separates Qwen3.8 Max exact-source benchmark rows from its aggregate ranking; weights, providers, harnesses, samples, and dates differ, making it a verifiable ledger rather than a unified independent rerun.Qwen3.8 Max: Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting GuideTurn Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling GuideTurn Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” ReportsTurn Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Roleplay: Direction Following, Reasoning Time, and Preset FeedbackTurn Qwen3.8-Max Roleplay: Direction Following, Reasoning Time, and Preset Feedback into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.