Qwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.
Qwen official blog · Read evidenceQwen3.7 Max · Reviews and evidence
Which Qwen3.7 Max conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
BenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.
BenchLM.ai · Read evidenceOfox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.
Ofox AI · Read evidenceFull reviews and related reading
Selected evidence
Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization Experiment
Qwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.
- Source-specific observation
- The May 20, 2026 Qwen release compares several models through Claude Code, OpenClaw, Qwen Code, and custom tools.
- Published conditions
- Not all prompts, seeds, failures, or runnable artifacts are public; the 35-hour kernel case uses the Zhenwu M890 PPU and a custom verifier.
Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed Ledger
BenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The ledger is current through August 17, 2026, with 437 tracked rows, 35 source-displayable rows, and a May 16 release label.
- Published conditions
- It has no common raw prompts, sampling settings, or complete traces, and does not independently verify API price or model ID.
Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real Tasks
Ofox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.
Unverified: the original source could not be rechecked.
- Source-specific observation
- Published June 2, updated August 3, 2026; tasks include a 1,200-line Python refactor, screenshot plus stack-trace debugging, and a 1,000-step Postgres migration.
- Published conditions
- Each model runs five times at temperature 0.2; the long-horizon task is unattended for four hours and quality is scored 1–5 by a senior reviewer.
Qwen3.7-Max: Artificial Analysis Intelligence Index, Cost, and Speed Benchmark
Artificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The page uses a May 20, 2026 release label, August 2026 data, and a 182-model comparison pool with first-party API inputs.
- Published conditions
- Cost per task is weighted by Intelligence Index and speed is tokens/s; exact prompts, repeats, and snapshots require the methods page for reproduction.
All sources
All sources
Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization Experiment
Qwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.
- Source-specific observation
- The May 20, 2026 Qwen release compares several models through Claude Code, OpenClaw, Qwen Code, and custom tools.
- Published conditions
- Not all prompts, seeds, failures, or runnable artifacts are public; the 35-hour kernel case uses the Zhenwu M890 PPU and a custom verifier.
Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed Ledger
BenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The ledger is current through August 17, 2026, with 437 tracked rows, 35 source-displayable rows, and a May 16 release label.
- Published conditions
- It has no common raw prompts, sampling settings, or complete traces, and does not independently verify API price or model ID.
Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real Tasks
Ofox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.
Unverified: the original source could not be rechecked.
- Source-specific observation
- Published June 2, updated August 3, 2026; tasks include a 1,200-line Python refactor, screenshot plus stack-trace debugging, and a 1,000-step Postgres migration.
- Published conditions
- Each model runs five times at temperature 0.2; the long-horizon task is unattended for four hours and quality is scored 1–5 by a senior reviewer.
Qwen3.7-Max: Artificial Analysis Intelligence Index, Cost, and Speed Benchmark
Artificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The page uses a May 20, 2026 release label, August 2026 data, and a 182-model comparison pool with first-party API inputs.
- Published conditions
- Cost per task is weighted by Intelligence Index and speed is tokens/s; exact prompts, repeats, and snapshots require the methods page for reproduction.
Qwen3.7-Max in Real-World Coding Tasks: Negative Instruction Confusion and Token Burn Benchmark
A Reddit developer report describes a Qwen3.7 Max “disable, do not delete” failure and 8.23M input tokens across three instructions and 134 internal interactions; client caching may confound it.
Unverified: the original source could not be rechecked.
- Source-specific observation
- Collected May 28 and summarized through July 15, 2026; the cases come from OpenCode coding tasks compared with several models.
- Published conditions
- It reports one semantic failure and one 8.23M-token run without controlled repeats, a common API harness, or complete logs.
Qwen3.7-Max: Arena.ai Blind Test Leaderboard and Domain Rankings
An ArenaAI community blind-test post places Qwen3.7 Max in text and domain rankings; results depend on the snapshot and voting pool rather than a fixed task suite.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The May 18, 2026 X/Arena post describes large-scale user-prompt pairwise voting across text, math, and professional domains.
- Published conditions
- It does not publish the full vote sample, snapshot, prompts, judging script, or repeats; preview rankings can move.
Qwen3.7-Max ITBench-AA Enterprise IT Operations and SRE Root-Cause Analysis Benchmark
ITBench-AA evaluates Qwen3.7 Max on offline incident snapshots containing Prometheus, OpenTelemetry, Kubernetes, logs, and topology evidence in the Stirrup sandbox; tools and payload shape the result.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The May 28, 2026 IBM/Artificial Analysis evaluation uses offline incident snapshots with metrics, traces, Kubernetes events, logs, and service topology.
- Published conditions
- The model gets shell and file-retrieval access in Stirrup; the result is not a tool-free chat score and full traces and deployment cost are not public.
Qwen3.7-Max: AA-Omniscience Knowledge Reliability and Hallucination Rate Benchmark
AA-Omniscience tests Qwen3.7 Max for knowledge uncertainty and hallucination, with an explicit abstention/not-attempted definition; it is a knowledge reliability benchmark, not a general agent test.
Unverified: the original source could not be rechecked.
- Source-specific observation
- Artificial Analysis runs standard knowledge prompts through official APIs on dedicated evaluation hardware, with data updated through August 2026.
- Published conditions
- Hallucination Rate is Incorrect/(Incorrect + Partial + Not Attempted); no extra deception condition is injected and full sample/repeat details require the methods page.
Qwen3.7 Max
Compare Qwen3.7 Max in Tabbit
Model access, features, and permissions depend on your current client account.