Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.7 Max · Media / benchmark · Independent measurement

Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed Ledger

BenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
The ledger is current through August 17, 2026, with 437 tracked rows, 35 source-displayable rows, and a May 16 release label.
Published conditions
It has no common raw prompts, sampling settings, or complete traces, and does not independently verify API price or model ID.

Key data and applicable tasks

One-sentence takeaway

BenchLM rates Qwen3.7 Max at 71.6/100, with a public rank of #16/218 and an evidence-verified rank of #13/104; it ranks #1 in multilingual performance but only #103 in Agentic, while its API price and model ID have not been independently verified.

Test environment

  • Dataset: BenchLM's public model directory and 437 tracked benchmark rows.

  • Timing: The page is marked “data as of 2026-08-17,” while the model page lists a 2026-05-16 release.

  • Evidence method: A Verified row must be tied to a public source; missing fields remain not sourced/not published and are not filled with estimates.

  • Runtime fields: The page records 204 tok/s and 14.07s to first token, but marks the direct context source as not stored; this cannot be treated as a complete runtime benchmark.

Inputs/configuration

The BenchLM page provides a summary of 35 source-displayable rows and some provider-exact values, but does not publish a unified original prompt, sampling parameters, or all runtime traces. Its “official exact-value snapshot” cites the Qwen release page dated 2026-05-16.

Results data

Overall and category results

ItemPage result
BenchLM score71.6/100
Public rank#16/218
Verified rank#13/104
Agentic39.8, #103/131
Coding52.8, #52/135
Reasoning85.8, unranked
Knowledge71.7, #22/57
Math81.8, unranked
Multilingual100.0, #1/13
Instruction Following92.6, #10/40
Speed204 tok/s; TTFT 14.07s
Context1M tokens

Coding results in the page ledger

  • LiveCodeBench 91.6%; listed by the page as best verified.

  • SWE-bench Verified 80.4%, 15.6 percentage points below the page's best-verified Claude Opus 5 at 96%.

  • SciCode 53.5%, 6.6 percentage points below best-verified Sakana Fugu at 60.1%.

  • SWE-bench Pro 60.6%, 19.7 percentage points below best-verified Claude Mythos 5 at 80.3%.

  • SWE Multilingual 78.3%, NL2Repo 47.2%, and Terminal-Bench 2.0 69.7%.

Conclusions

  • Multilingual performance and instruction following are the most prominent relative strengths in BenchLM's evidence; coding evidence is fairly complete, but not every SWE/Terminal metric is near the top of the leaderboard.

  • Agentic #103/131 is in clear tension with the official long-horizon case, showing that a “35-hour demo” cannot substitute for independent Agent metrics across tasks and harnesses.

  • The 204 tok/s speed figure is only a directory runtime entry; with 14.07s TTFT and variables such as region, concurrency, and long context, it must be retested with your own API trace.

  • BenchLM marks the model as superseded, and Qwen3.8 Max Preview has appeared in the lineage; production selection should first confirm whether the stable 3.7 snapshot is still required.

Limitations

  • BenchLM is a third-party aggregator; its composite score includes disclosed weighting and missing-data handling. Do not treat 71.6 as a single benchmark.

  • API model ID, price, maximum output, knowledge cutoff, modality, and other fields are marked as unpublished or not independently verified.

  • Some benchmarks are provider exact or display only, and their harnesses and evaluators are not fully consistent.

  • The direct sources for the speed and context fields are incomplete; they cannot replace the official API pricing page or your own load test.

Reproduction steps

  1. Save the BenchLM collection date, model lifecycle status, and evidence label for each ledger entry.

  2. Using the official Qwen page, pin a qwen3.7-max dated snapshot and run LiveCodeBench, SWE, and multilingual tasks separately.

  3. For Agent tasks, additionally record TTFT, output speed, tool calls, context length, failures/retries, and cost-per-success.

  4. Report “public rank,” “verified rank,” category scores, and speed separately; do not recompute them into an unexplained total score.

  5. Run an A/B test against Qwen3.8 Max Preview with the same harness to assess the reproducibility and supply risks of continuing with 3.7.

Source excerpt or observation (brief excerpt for compliance only)

BenchLM's decision prompt is “Validate before choosing,” because the 35 visible results still leave some benchmark slots empty.

What this supports

  • It supports using category coverage and speed fields to choose local reproduction tests.

What this does not support

  • It cannot support a uniform leaderboard, current cost guarantee, or provider-neutral latency claim.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

BenchLM.ai · BenchLM · Original publication date 2026-05-16 · Site edit date 2026-09-20

Open original source

Qwen3.7 Max

Compare Qwen3.7 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.7 Max: What It Is, Access Routes, and Where It Fits

A sourced Qwen3.7 Max overview covering the dated snapshot, one-million-token API boundary, agentic use cases, pricing separation, and practical risks.

Related reviews

Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization ExperimentQwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real TasksOfox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.Qwen3.7-Max: Artificial Analysis Intelligence Index, Cost, and Speed BenchmarkArtificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.Qwen3.7-Max ITBench-AA Enterprise IT Operations and SRE Root-Cause Analysis BenchmarkITBench-AA evaluates Qwen3.7 Max on offline incident snapshots containing Prometheus, OpenTelemetry, Kubernetes, logs, and topology evidence in the Stirrup sandbox; tools and payload shape the result.Qwen3.7-Max: Long-Horizon Agents, Frontend Prototypes, and Office PromptsQwen’s long-horizon case separates frontend, office, and multi-step agent work into verifiable stages, with tools, timeouts, and final acceptance recorded separately.Qwen3.7-Max: Alibaba Cloud Model Studio Versions, Pricing, and Cache ConfigurationThe Model Studio page separates Qwen3.7 Max snapshots, the million-token context, cache billing, and regional limits so a test can fix version and cost assumptions first.Qwen3.7-Max: Three.js Electronic Rubik's Cube and 3D Physics Interaction Prototype PromptAlibaba Cloud’s public Three.js case turns visual references, interaction rules, and physics checks into an electronic cube prototype task for browser-based visual prototyping.Qwen3.7-Max: Long-Horizon Agent Prompts and Acceptance Closed Loop for GPU Kernel OptimizationThe Qwen team’s GPU-kernel case connects performance hypotheses, compilation tests, benchmark regressions, and long-horizon progress logs; hardware and verifier scope are critical boundaries.