Artificial Analysis's current page gives Qwen3.8 Max an Intelligence Index score of 58 (No. 10 among 180 comparable models), while also recording its low speed of about 45 token/s, roughly 150M total output tokens across the index evaluations, and an approximate per-task cost of $1.13. It therefore looks more like a “high-quality but slow and highly deliberative” Agent model than a low-latency chat model.
Suitable tasks: Agent tasks requiring tool calls, long-chain delivery, code execution, and high completion quality; model selection that compares models by task cost rather than unit price alone.
Unsuitable tasks: Highly real-time conversations, high-throughput services with short answers, and scenarios that treat a text-based English composite index as a direct proxy for multimodal or Chinese-language business quality.
Applicable model version: The reasoning version of Qwen3.8 Max; page data may change with the provider, evaluation version, and date.
Applicable client, Agent, or API: Artificial Analysis's unified evaluation environment; this does not equal actual latency in Qwen Studio, DashScope, or third-party routers.
Recommended reasoning tier and parameters: The complete prompt for a reproducible single-model test is not public; the methodology page generally uses temperature 0.6 for reasoning models and the maximum output limit provided by the model. Before launch, test low, medium, and high reasoning tiers separately in your own Agent harness.
Artificial Analysis Intelligence Index v4.1.1 includes GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, Humanity’s Last Exam, GPQA Diamond, and CritPt, for nine evaluations in total.
The index is weighted as Agents 34%, Coding 24%, Scientific Reasoning 24%, and General 18%; different evaluations use different numbers of repetitions and tool configurations.
The public methodology page gives sample sizes including: GDPval-AA v2, 220 tasks, 1 run; τ³-Banking, 97 questions, 5 runs; Terminal-Bench v2.1, 89 tasks, 3 runs; SciCode, 288 subquestions, 3 runs; AA-LCR, 100 questions, 3 runs; AA-Omniscience, 6,000 questions, 1 run; HLE, 2,158 questions, 1 run; GPQA Diamond, 198 questions, 5 runs; and CritPt, 70 questions, 5 runs.
The methodology page states that the evaluations use zero-shot instructions, unified scoring, and pass@1; GDPval-AA v2 uses the Stirrup harness, an E2B sandbox, file output, and web/code tools.
For review, open both the model page and the Artificial Analysis Intelligence Index methodology page, and record the page date, evaluation version, price, and model endpoint; do not screenshot only the composite score.
Overall quality: 58; the model summary on the page shows No. 10 among 180 comparable models.
Cost: $2.00 / 1M tokens input, $6.00 / 1M tokens output, 88% cache discount; the Artificial Analysis page shows an approximate Intelligence Index per-task cost of $1.13 and a full-index evaluation cost of approximately $1,741.41.
Speed: Output speed of about 44.6 token/s, listed by the page as No. 117/180 among comparable models; the summary places it in speed tier 1/4.
Output volume: Approximately 150M tokens generated cumulatively across the index evaluations; the page places it at No. 70/180 among comparable models for verbosity and explicitly warns that the model is very verbose.
Input capabilities: Supports text, image, and video input, with text output; the context window is listed as 1M tokens.
Methodological boundary: The methodology page describes the Intelligence Index as an English, text-only evaluation; visual, audio, and multilingual capabilities are measured separately and are not included in the composite score.
This source supports three conclusions that hold simultaneously: Qwen3.8-Max's overall capability is near the top; its task cost is not determined only by the $2/$6 unit prices, but is amplified by reasoning and output volume; and in human use, waiting time and answer length may become bottlenecks before the overall score does. When deploying it as the primary model for high-value Agents, treat “successful cost per task, completion latency, number of tool calls, and output tokens” as one acceptance sheet rather than looking only at the Intelligence Index.
The page is a dynamic leaderboard; rankings, index versions, providers, and prices may change, so data collected on one day does not represent a permanent ranking.
The Intelligence Index is a weighted aggregate score focused on English text and cannot directly represent image, video, Chinese-language, or domain-specific tasks.
Artificial Analysis's methodology page publishes the evaluation structure and parameters, but the complete private inputs, execution traces, and details of every model endpoint are not all shown on the model page; readers cannot fully reproduce a score of 58 from the page alone.
Per-task cost includes weighted input, cache, reasoning, and answer tokens; it should not be treated as the quote for one ordinary API request.
Fix the collection date and save the model page's score, ranking, input/output prices, speed, context, and output-volume fields.
Fix Artificial Analysis Intelligence Index v4.1.1 and the weights of its nine evaluations; do not mix data from later versions into the same table.
Build a zero-shot regression set of the same tasks on your own Qwen3.8-Max endpoint, recording at least success rate, time to first token, total elapsed time, reasoning tokens, answer tokens, tool-call count, and cost per task.
Rerun with low/medium/high reasoning effort, compare “quality improvement ÷ additional tokens” and “successful-task cost,” and do not use this to claim that you reproduced Artificial Analysis's absolute score.
The page summary describes Qwen3.8 Max as “notably slow and very verbose”; the methodology page also states that the Intelligence Index is only a composite comparison metric and cannot be applied directly to every use case. The two points should be cited together.
Qwen3.8 Max