In the Arena.ai double-blind battle rankings, Qwen3.7-Max ranked #13 globally on the overall Text Arena leaderboard, Alibaba ranked #6 among global AI labs, and the model secured top-10 positions across specialized sub-leaderboards including Math (#7), Expert (#9), Software & IT (#9), and Coding (#10).
Evaluation platform: Arena.ai (based on real-user anonymous blind test battles and an Elo dynamic rating system).
Sample scope: Millions of real-world prompt battles from developers and users worldwide, eliminating benchmark overfitting and prompt contamination bias.
Model evaluated: Qwen3.7-Max Preview (Text Arena).
Evaluation mode: Open-domain real user interaction prompts, covering coding, mathematics, long-form writing, complex multi-turn dialogues, and system instruction following.
Comparison pool: Includes OpenAI GPT series, Anthropic Claude series, Google Gemini series, and leading open-source derivative models.
| Category | Official rank | Tier positioning |
|---|---|---|
| Text Arena Overall | #13 | Top-tier global flagship model |
| Math | #7 | Substantially outperforms general LLMs of the same generation |
| Expert | #9 | Top 10 in high-difficulty academic and specialized professional consulting |
| Software & IT | #9 | Top 10 in enterprise system architecture and IT operations |
| Coding | #10 | Top 10 in core software engineering and bug fixing |
| Lab Rank | #6 | Alibaba ranks among the top 6 global AI labs |
Math (#7): Trails only leading pure-reasoning models (such as the OpenAI o-series / GPT-5.x reasoning flagships and the DeepSeek R-series), representing top-tier performance among general-purpose foundation models.
Software & IT (#9): Corroborates its performance on the ITBench-AA leaderboard, demonstrating solid contextual understanding when handling real-world systems operations, network policies, and cloud-native infrastructure diagnostics.
Coding (#10): Consistently ranks in the global top 10 in real-world blind testing, indicating that its code generation quality is directly validated by the broader developer community.
Blind test data effectively dispels concerns about overfitting to academic benchmarks: Qwen3.7-Max demonstrates well-rounded, high-level capabilities in real-user-driven Arena blind testing, showing standout competitiveness particularly in mathematics and IT operations tasks.
Its strong rankings in Math and IT (#7 and #9) align closely with positive feedback reported in downstream enterprise SRE, quantitative finance code review, and other vertical production scenarios.
Early Arena ratings were primarily based on a Preview snapshot; rankings remain dynamic as various model providers iterate rapidly.
Blind testing emphasizes perceived response quality in single-turn or short multi-turn interactions, offering limited coverage for long-horizon autonomous agents requiring external tool calling (such as 1,000+ step CLI migration tasks).
Log in to the Arena.ai / LMSYS platform and select qwen3.7-max in Direct Chat or Side-by-Side mode.
Construct multi-turn battle prompts covering mathematical proofs, algorithmic implementations, and complex system configuration diagnostics.
Track the model's win rate and Elo rating trajectory under anonymous blind evaluations.
Official Arena.ai announcement: “In Text Arena, Qwen3.7 Max Preview ranks #13 overall. Alibaba is now the #6 lab in this arena: #7 Math, #9 Expert, #9 Software & IT, #10 Coding”.
Qwen3.7 Max