Qwen3.7 Max · Media / benchmark · Independent measurement
Artificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
Artificial Analysis independent benchmark results show Qwen3.7-Max scoring 47 on the Intelligence Index (top 23%) , ranking #9 with an output speed of 206.4 tok/s, and achieving a per-task cost of $0.54—significantly lower than Opus 5 and GPT-5.6 Sol—though its tendency to generate 100M tokens reflects noticeable long-thought verbosity.
Evaluation framework: Artificial Analysis Intelligence Index v4.1.1.
Benchmark subsets: 9 controlled evaluations including GDPval-AA v2, tau^3-Banking (𝜏³-Banking) , Terminal-Bench v2.1, SciCode, Humanity's Last Exam (HLE) , GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR.
Inference environment: Tested on the official Alibaba Cloud Model Studio inference endpoint under a default 10k input token workload.
Comparison pool: Cross-model comparison across 182 frontier reasoning and proprietary models worldwide.
Evaluation mode: Full Reasoning / Thinking mode enabled.
Input/output modalities: Pure text (Text-only) , supporting up to a 1M token context per single request.
Pricing baseline: Input $2.50 / 1M tokens, Output $7.50 / 1M tokens, with an 80% context caching discount.
Total consumption: In completing the entire Intelligence Index benchmark suite, a cumulative 100M output tokens were generated, totaling $1,063.86 in testing expenses.
| Dimension | Measured value | Relative rank / Percentile | Benchmark median |
|---|---|---|---|
| Intelligence Index | 47 points | #42 / 182 | 35 points |
| Output Speed | 206.4 tok/s | #9 / 182 | 76 tok/s |
| Time to First Token (TTFT) | 2.27s | Better than median | 2.79s |
| Weighted Cost per Task | $0.54 | #52 / 182 | - |
| Evaluation Total Output Tokens (Verbosity) | 100M tokens | #56 / 182 (tending verbose) | 72M tokens |
| Model | Reasoning tier | Cost per task (USD) | Intelligence Index |
|---|---|---|---|
| GPT-5.6 Luna | max | $0.05 | 52 |
| DeepSeek V4 Pro (0813) | max | $0.25 | 53 |
| Gemini 3.7 Flash | high | $0.40 | 56 |
| Qwen3.7-Max | reasoning | $0.54 | 47 |
| GLM-5.3 | max | $0.68 | 60 |
| Grok 4.6 | high | $0.84 | 61 |
| Kimi K3 | max | $0.84 | 60 |
| GPT-5.6 Sol | max | $1.23 | 61 |
| Claude Opus 5 | max | $2.34 | 63 |
| Claude Fable 5 | fallback | $3.14 | 62 |
| Model | Output speed (tok/s) |
|---|---|
| Gemini 3.7 Flash (high) | 365 |
| Qwen3.7-Max | 206 |
| GPT-5.6 Luna (max) | 149 |
| Nemotron 3 Ultra | 126 |
| GLM-5.3 (max) | 93 |
| DeepSeek V4 Pro (0813) | 80 |
| Claude Fable 5 | 75 |
| Claude Opus 5 (max) | 59 |
| Kimi K3 (max) | 39 |
Prominent speed and throughput advantages: While sustaining a 47-point intelligence level, Qwen3.7-Max delivers an output rate of 206.4 tok/s—3.5x faster than Claude Opus 5 (59 tok/s) and 5.3x faster than Kimi K3 (39 tok/s) —making it exceptionally well-suited for interactive coding Agents.
Competitive cost-effectiveness: At $0.54 per task, its cost is only 23% of Claude Opus 5 ($2.34) , offering a clear cost advantage in enterprise IT and code refactoring scenarios requiring frequent reasoning calls.
Elevated reasoning verbosity: Generating 100M tokens during the evaluation is noticeably higher than the industry median of 72M, indicating that the model tends to expand into very long chains of thought (Long-thought) during complex logical deduction, requiring appropriate max_tokens budget management.
Text-only input limitation: The model does not support visual or screenshot inputs (Multimodal Image Drop) , making it unusable for direct frontend UI screenshot reviews.
Deprecation and migration: Following the release of Qwen3.8-Max (56 points) , the vendor has categorized Qwen3.7-Max under Deprecated archive-tracking status; long-term production deployments should evaluate a smooth migration path to version 3.8.
Call the qwen3.7-max-2026-05-20 snapshot via the official Alibaba Cloud Model Studio endpoint with enable_thinking: true enabled.
Run standard 10k input token baseline requests, recording TTFT (expected within the 2.0s–2.5s range) and streaming output tok/s.
Profile the proportion of thinking tokens in long-horizon reasoning tasks (Thinking vs. Final Answer) and calculate the actual completion cost per task.
Artificial Analysis benchmark summary: “Qwen3.7 Max is amongst the leading models in intelligence... It's also notably fast, however somewhat verbose. At 206 tokens per second, Qwen3.7 Max is notably fast”.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Artificial Analysis · Artificial Analysis Team · Original publication date 2026-05-20 · Site edit date 2026-09-20
Open original sourceQwen3.7 Max
Download the Tabbit client to check model access