Qwen3.7 Max

Qwen3.7 Max · Reviews and evidence

Which Qwen3.7 Max conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

Qwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.

Qwen official blog · Read evidence

BenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.

BenchLM.ai · Read evidence

Ofox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.

Ofox AI · Read evidence

Full reviews and related reading

Read the full analysis

Overview · English

Qwen3.7 Max: What It Is, Access Routes, and Where It Fits

A sourced Qwen3.7 Max overview covering the dated snapshot, one-million-token API boundary, agentic use cases, pricing separation, and practical risks.

Selected evidence

Media / benchmarkVendor report

Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization Experiment

Qwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.

SourceQwen official blog
Published2026-05-20
Collected2026-09-20
Source-specific observation
The May 20, 2026 Qwen release compares several models through Claude Code, OpenClaw, Qwen Code, and custom tools.
Published conditions
Not all prompts, seeds, failures, or runnable artifacts are public; the 35-hour kernel case uses the Zhenwu M890 PPU and a custom verifier.
Capability
Media / benchmarkIndependent measurement

Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed Ledger

BenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.

SourceBenchLM.ai
Published2026-05-16
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
The ledger is current through August 17, 2026, with 437 tracked rows, 35 source-displayable rows, and a May 16 release label.
Published conditions
It has no common raw prompts, sampling settings, or complete traces, and does not independently verify API price or model ID.
Capability
Media / benchmarkIndependent measurement

Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real Tasks

Ofox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.

SourceOfox AI
Published2026-06-02
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
Published June 2, updated August 3, 2026; tasks include a 1,200-line Python refactor, screenshot plus stack-trace debugging, and a 1,000-step Postgres migration.
Published conditions
Each model runs five times at temperature 0.2; the long-horizon task is unattended for four hours and quality is scored 1–5 by a senior reviewer.
Capability
Media / benchmarkIndependent measurement

Qwen3.7-Max: Artificial Analysis Intelligence Index, Cost, and Speed Benchmark

Artificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.

SourceArtificial Analysis
Published2026-05-20
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
The page uses a May 20, 2026 release label, August 2026 data, and a 182-model comparison pool with first-party API inputs.
Published conditions
Cost per task is weighted by Intelligence Index and speed is tokens/s; exact prompts, repeats, and snapshots require the methods page for reproduction.
Capability

All sources

All sources

8 / 8
Media / benchmarkVendor report

Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization Experiment

Qwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.

SourceQwen official blog
Published2026-05-20
Collected2026-09-20
Source-specific observation
The May 20, 2026 Qwen release compares several models through Claude Code, OpenClaw, Qwen Code, and custom tools.
Published conditions
Not all prompts, seeds, failures, or runnable artifacts are public; the 35-hour kernel case uses the Zhenwu M890 PPU and a custom verifier.
Capability
Media / benchmarkIndependent measurement

Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed Ledger

BenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.

SourceBenchLM.ai
Published2026-05-16
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
The ledger is current through August 17, 2026, with 437 tracked rows, 35 source-displayable rows, and a May 16 release label.
Published conditions
It has no common raw prompts, sampling settings, or complete traces, and does not independently verify API price or model ID.
Capability
Media / benchmarkIndependent measurement

Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real Tasks

Ofox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.

SourceOfox AI
Published2026-06-02
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
Published June 2, updated August 3, 2026; tasks include a 1,200-line Python refactor, screenshot plus stack-trace debugging, and a 1,000-step Postgres migration.
Published conditions
Each model runs five times at temperature 0.2; the long-horizon task is unattended for four hours and quality is scored 1–5 by a senior reviewer.
Capability
Media / benchmarkIndependent measurement

Qwen3.7-Max: Artificial Analysis Intelligence Index, Cost, and Speed Benchmark

Artificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.

SourceArtificial Analysis
Published2026-05-20
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
The page uses a May 20, 2026 release label, August 2026 data, and a 182-model comparison pool with first-party API inputs.
Published conditions
Cost per task is weighted by Intelligence Index and speed is tokens/s; exact prompts, repeats, and snapshots require the methods page for reproduction.
Capability
CommunityPersonal experience

Qwen3.7-Max in Real-World Coding Tasks: Negative Instruction Confusion and Token Burn Benchmark

A Reddit developer report describes a Qwen3.7 Max “disable, do not delete” failure and 8.23M input tokens across three instructions and 134 internal interactions; client caching may confound it.

SourceReddit (r/opencode, r/QwenAI, r/LocalLLaMA)
Published2026-05-28
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
Collected May 28 and summarized through July 15, 2026; the cases come from OpenCode coding tasks compared with several models.
Published conditions
It reports one semantic failure and one 8.23M-token run without controlled repeats, a common API harness, or complete logs.
Capability
CommunityPersonal experience

Qwen3.7-Max: Arena.ai Blind Test Leaderboard and Domain Rankings

An ArenaAI community blind-test post places Qwen3.7 Max in text and domain rankings; results depend on the snapshot and voting pool rather than a fixed task suite.

SourceArena.ai (LMSYS Chatbot Arena)
Published2026-05-18
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
The May 18, 2026 X/Arena post describes large-scale user-prompt pairwise voting across text, math, and professional domains.
Published conditions
It does not publish the full vote sample, snapshot, prompts, judging script, or repeats; preview rankings can move.
Capability
Media / benchmarkIndependent measurement

Qwen3.7-Max ITBench-AA Enterprise IT Operations and SRE Root-Cause Analysis Benchmark

ITBench-AA evaluates Qwen3.7 Max on offline incident snapshots containing Prometheus, OpenTelemetry, Kubernetes, logs, and topology evidence in the Stirrup sandbox; tools and payload shape the result.

SourceArtificial Analysis & IBM Research
Published2026-05-28
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
The May 28, 2026 IBM/Artificial Analysis evaluation uses offline incident snapshots with metrics, traces, Kubernetes events, logs, and service topology.
Published conditions
The model gets shell and file-retrieval access in Stirrup; the result is not a tool-free chat score and full traces and deployment cost are not public.
Capability
Media / benchmarkIndependent measurement

Qwen3.7-Max: AA-Omniscience Knowledge Reliability and Hallucination Rate Benchmark

AA-Omniscience tests Qwen3.7 Max for knowledge uncertainty and hallucination, with an explicit abstention/not-attempted definition; it is a knowledge reliability benchmark, not a general agent test.

SourceArtificial Analysis
Published2026-05-20
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
Artificial Analysis runs standard knowledge prompts through official APIs on dedicated evaluation hardware, with data updated through August 2026.
Published conditions
Hallucination Rate is Incorrect/(Incorrect + Partial + Not Attempted); no extra deception condition is injected and full sample/repeat details require the methods page.
Capability

Qwen3.7 Max

Compare Qwen3.7 Max in Tabbit

Model access, features, and permissions depend on your current client account.