Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.7 Max · Media / benchmark · Independent measurement

Qwen3.7-Max: Artificial Analysis Intelligence Index, Cost, and Speed Benchmark

Artificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
The page uses a May 20, 2026 release label, August 2026 data, and a 182-model comparison pool with first-party API inputs.
Published conditions
Cost per task is weighted by Intelligence Index and speed is tokens/s; exact prompts, repeats, and snapshots require the methods page for reproduction.

Key data and applicable tasks

One-sentence takeaway

Artificial Analysis independent benchmark results show Qwen3.7-Max scoring 47 on the Intelligence Index (top 23%) , ranking #9 with an output speed of 206.4 tok/s, and achieving a per-task cost of $0.54—significantly lower than Opus 5 and GPT-5.6 Sol—though its tendency to generate 100M tokens reflects noticeable long-thought verbosity.

Test environment

  • Evaluation framework: Artificial Analysis Intelligence Index v4.1.1.

  • Benchmark subsets: 9 controlled evaluations including GDPval-AA v2, tau^3-Banking (𝜏³-Banking) , Terminal-Bench v2.1, SciCode, Humanity's Last Exam (HLE) , GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR.

  • Inference environment: Tested on the official Alibaba Cloud Model Studio inference endpoint under a default 10k input token workload.

  • Comparison pool: Cross-model comparison across 182 frontier reasoning and proprietary models worldwide.

Inputs/configuration

  • Evaluation mode: Full Reasoning / Thinking mode enabled.

  • Input/output modalities: Pure text (Text-only) , supporting up to a 1M token context per single request.

  • Pricing baseline: Input $2.50 / 1M tokens, Output $7.50 / 1M tokens, with an 80% context caching discount.

  • Total consumption: In completing the entire Intelligence Index benchmark suite, a cumulative 100M output tokens were generated, totaling $1,063.86 in testing expenses.

Results data

Core capabilities and performance rankings

DimensionMeasured valueRelative rank / PercentileBenchmark median
Intelligence Index47 points#42 / 18235 points
Output Speed206.4 tok/s#9 / 18276 tok/s
Time to First Token (TTFT)2.27sBetter than median2.79s
Weighted Cost per Task$0.54#52 / 182-
Evaluation Total Output Tokens (Verbosity)100M tokens#56 / 182 (tending verbose)72M tokens

Cross-model weighted cost per task comparison (Cost per Task)

ModelReasoning tierCost per task (USD)Intelligence Index
GPT-5.6 Lunamax$0.0552
DeepSeek V4 Pro (0813)max$0.2553
Gemini 3.7 Flashhigh$0.4056
Qwen3.7-Maxreasoning$0.5447
GLM-5.3max$0.6860
Grok 4.6high$0.8461
Kimi K3max$0.8460
GPT-5.6 Solmax$1.2361
Claude Opus 5max$2.3463
Claude Fable 5fallback$3.1462

Cross-model output speed comparison (Output Tokens / sec)

ModelOutput speed (tok/s)
Gemini 3.7 Flash (high)365
Qwen3.7-Max206
GPT-5.6 Luna (max)149
Nemotron 3 Ultra126
GLM-5.3 (max)93
DeepSeek V4 Pro (0813)80
Claude Fable 575
Claude Opus 5 (max)59
Kimi K3 (max)39

Conclusions

  • Prominent speed and throughput advantages: While sustaining a 47-point intelligence level, Qwen3.7-Max delivers an output rate of 206.4 tok/s—3.5x faster than Claude Opus 5 (59 tok/s) and 5.3x faster than Kimi K3 (39 tok/s) —making it exceptionally well-suited for interactive coding Agents.

  • Competitive cost-effectiveness: At $0.54 per task, its cost is only 23% of Claude Opus 5 ($2.34) , offering a clear cost advantage in enterprise IT and code refactoring scenarios requiring frequent reasoning calls.

  • Elevated reasoning verbosity: Generating 100M tokens during the evaluation is noticeably higher than the industry median of 72M, indicating that the model tends to expand into very long chains of thought (Long-thought) during complex logical deduction, requiring appropriate max_tokens budget management.

Limitations

  • Text-only input limitation: The model does not support visual or screenshot inputs (Multimodal Image Drop) , making it unusable for direct frontend UI screenshot reviews.

  • Deprecation and migration: Following the release of Qwen3.8-Max (56 points) , the vendor has categorized Qwen3.7-Max under Deprecated archive-tracking status; long-term production deployments should evaluate a smooth migration path to version 3.8.

Reproduction steps

  1. Call the qwen3.7-max-2026-05-20 snapshot via the official Alibaba Cloud Model Studio endpoint with enable_thinking: true enabled.

  2. Run standard 10k input token baseline requests, recording TTFT (expected within the 2.0s–2.5s range) and streaming output tok/s.

  3. Profile the proportion of thinking tokens in long-horizon reasoning tasks (Thinking vs. Final Answer) and calculate the actual completion cost per task.

Source excerpt or observation (brief excerpt for compliance only)

Artificial Analysis benchmark summary: “Qwen3.7 Max is amongst the leading models in intelligence... It's also notably fast, however somewhat verbose. At 206 tokens per second, Qwen3.7 Max is notably fast”.

What this supports

  • It supports comparing the displayed cost-speed-intelligence snapshot and selecting a reasoning budget to reproduce.

What this does not support

  • It cannot be treated as a fixed price, universal quality ranking, or stable behavior after a provider update.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis Team · Original publication date 2026-05-20 · Site edit date 2026-09-20

Open original source

Qwen3.7 Max

Compare Qwen3.7 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.7 Max: What It Is, Access Routes, and Where It Fits

A sourced Qwen3.7 Max overview covering the dated snapshot, one-million-token API boundary, agentic use cases, pricing separation, and practical risks.

Related reviews

Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization ExperimentQwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed LedgerBenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real TasksOfox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.Qwen3.7-Max ITBench-AA Enterprise IT Operations and SRE Root-Cause Analysis BenchmarkITBench-AA evaluates Qwen3.7 Max on offline incident snapshots containing Prometheus, OpenTelemetry, Kubernetes, logs, and topology evidence in the Stirrup sandbox; tools and payload shape the result.Qwen3.7-Max: Long-Horizon Agents, Frontend Prototypes, and Office PromptsQwen’s long-horizon case separates frontend, office, and multi-step agent work into verifiable stages, with tools, timeouts, and final acceptance recorded separately.Qwen3.7-Max: Alibaba Cloud Model Studio Versions, Pricing, and Cache ConfigurationThe Model Studio page separates Qwen3.7 Max snapshots, the million-token context, cache billing, and regional limits so a test can fix version and cost assumptions first.Qwen3.7-Max: Three.js Electronic Rubik's Cube and 3D Physics Interaction Prototype PromptAlibaba Cloud’s public Three.js case turns visual references, interaction rules, and physics checks into an electronic cube prototype task for browser-based visual prototyping.Qwen3.7-Max: Long-Horizon Agent Prompts and Acceptance Closed Loop for GPU Kernel OptimizationThe Qwen team’s GPU-kernel case connects performance hypotheses, compilation tests, benchmark regressions, and long-horizon progress logs; hardware and verifier scope are critical boundaries.