Artificial Analysis independent benchmark results show Qwen3.7-Max scoring 47 on the Intelligence Index (top 23%) , ranking #9 with an output speed of 206.4 tok/s, and achieving a per-task cost of $0.54—significantly lower than Opus 5 and GPT-5.6 Sol—though its tendency to generate 100M tokens reflects noticeable long-thought verbosity.
Evaluation framework: Artificial Analysis Intelligence Index v4.1.1.
Benchmark subsets: 9 controlled evaluations including GDPval-AA v2, tau^3-Banking (𝜏³-Banking) , Terminal-Bench v2.1, SciCode, Humanity's Last Exam (HLE) , GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR.
Inference environment: Tested on the official Alibaba Cloud Model Studio inference endpoint under a default 10k input token workload.
Comparison pool: Cross-model comparison across 182 frontier reasoning and proprietary models worldwide.
Evaluation mode: Full Reasoning / Thinking mode enabled.
Input/output modalities: Pure text (Text-only) , supporting up to a 1M token context per single request.
Pricing baseline: Input $2.50 / 1M tokens, Output $7.50 / 1M tokens, with an 80% context caching discount.
Total consumption: In completing the entire Intelligence Index benchmark suite, a cumulative 100M output tokens were generated, totaling $1,063.86 in testing expenses.
| Dimension | Measured value | Relative rank / Percentile | Benchmark median |
|---|---|---|---|
| Intelligence Index | 47 points | #42 / 182 | 35 points |
| Output Speed | 206.4 tok/s | #9 / 182 | 76 tok/s |
| Time to First Token (TTFT) | 2.27s | Better than median | 2.79s |
| Weighted Cost per Task | $0.54 | #52 / 182 | - |
| Evaluation Total Output Tokens (Verbosity) | 100M tokens | #56 / 182 (tending verbose) | 72M tokens |
| Model | Reasoning tier | Cost per task (USD) | Intelligence Index |
|---|---|---|---|
| GPT-5.6 Luna | max | $0.05 | 52 |
| DeepSeek V4 Pro (0813) | max | $0.25 | 53 |
| Gemini 3.7 Flash | high | $0.40 | 56 |
| Qwen3.7-Max | reasoning | $0.54 | 47 |
| GLM-5.3 | max | $0.68 | 60 |
| Grok 4.6 | high | $0.84 | 61 |
| Kimi K3 | max | $0.84 | 60 |
| GPT-5.6 Sol | max | $1.23 | 61 |
| Claude Opus 5 | max | $2.34 | 63 |
| Claude Fable 5 | fallback | $3.14 | 62 |
| Model | Output speed (tok/s) |
|---|---|
| Gemini 3.7 Flash (high) | 365 |
| Qwen3.7-Max | 206 |
| GPT-5.6 Luna (max) | 149 |
| Nemotron 3 Ultra | 126 |
| GLM-5.3 (max) | 93 |
| DeepSeek V4 Pro (0813) | 80 |
| Claude Fable 5 | 75 |
| Claude Opus 5 (max) | 59 |
| Kimi K3 (max) | 39 |
Prominent speed and throughput advantages: While sustaining a 47-point intelligence level, Qwen3.7-Max delivers an output rate of 206.4 tok/s—3.5x faster than Claude Opus 5 (59 tok/s) and 5.3x faster than Kimi K3 (39 tok/s) —making it exceptionally well-suited for interactive coding Agents.
Competitive cost-effectiveness: At $0.54 per task, its cost is only 23% of Claude Opus 5 ($2.34) , offering a clear cost advantage in enterprise IT and code refactoring scenarios requiring frequent reasoning calls.
Elevated reasoning verbosity: Generating 100M tokens during the evaluation is noticeably higher than the industry median of 72M, indicating that the model tends to expand into very long chains of thought (Long-thought) during complex logical deduction, requiring appropriate max_tokens budget management.
Text-only input limitation: The model does not support visual or screenshot inputs (Multimodal Image Drop) , making it unusable for direct frontend UI screenshot reviews.
Deprecation and migration: Following the release of Qwen3.8-Max (56 points) , the vendor has categorized Qwen3.7-Max under Deprecated archive-tracking status; long-term production deployments should evaluate a smooth migration path to version 3.8.
Call the qwen3.7-max-2026-05-20 snapshot via the official Alibaba Cloud Model Studio endpoint with enable_thinking: true enabled.
Run standard 10k input token baseline requests, recording TTFT (expected within the 2.0s–2.5s range) and streaming output tok/s.
Profile the proportion of thinking tokens in long-horizon reasoning tasks (Thinking vs. Final Answer) and calculate the actual completion cost per task.
Artificial Analysis benchmark summary: “Qwen3.7 Max is amongst the leading models in intelligence... It's also notably fast, however somewhat verbose. At 206 tokens per second, Qwen3.7 Max is notably fast”.
Qwen3.7 Max