Qwen3.7-Max: Long-Horizon Agents, Frontend Prototypes, and Office Prompts
Qwen’s long-horizon case separates frontend, office, and multi-step agent work into verifiable stages, with tools, timeouts, and final acceptance recorded separately.
Prepare
task goal, source or reference material, runtime constraints, acceptance criteria
Runtime
Qwen3.7 Max client or API; pin the live model ID, provider, tools, permissions, and snapshot before execution.
Qwen3.7-Max: Alibaba Cloud Model Studio Versions, Pricing, and Cache Configuration
The Model Studio page separates Qwen3.7 Max snapshots, the million-token context, cache billing, and regional limits so a test can fix version and cost assumptions first.
Prepare
task goal, source or reference material, runtime constraints, acceptance criteria
Runtime
Qwen3.7 Max client or API; pin the live model ID, provider, tools, permissions, and snapshot before execution.
Qwen3.7-Max: Three.js Electronic Rubik's Cube and 3D Physics Interaction Prototype Prompt
Alibaba Cloud’s public Three.js case turns visual references, interaction rules, and physics checks into an electronic cube prototype task for browser-based visual prototyping.
Prepare
task goal, source or reference material, runtime constraints, acceptance criteria
Runtime
Qwen3.7 Max client or API; pin the live model ID, provider, tools, permissions, and snapshot before execution.
Qwen3.7-Max: Long-Horizon Agent Prompts and Acceptance Closed Loop for GPU Kernel Optimization
The Qwen team’s GPU-kernel case connects performance hypotheses, compilation tests, benchmark regressions, and long-horizon progress logs; hardware and verifier scope are critical boundaries.
Prepare
task goal, source or reference material, runtime constraints, acceptance criteria
Runtime
Qwen3.7 Max client or API; pin the live model ID, provider, tools, permissions, and snapshot before execution.
Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization Experiment
Qwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.
Evidence
Vendor report
Boundary
It cannot show that every agent harness can run unattended for 35 hours or that the scores transfer across hardware.
Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed Ledger
BenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.
Evidence
Independent measurement
Boundary
It cannot support a uniform leaderboard, current cost guarantee, or provider-neutral latency claim.
Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real Tasks
Ofox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.
Evidence
Independent measurement
Boundary
It cannot represent all repositories, longer autonomous runs, or current provider pricing outside those snapshots.