Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.7 Max · Media / benchmark · Independent measurement

Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real Tasks

Ofox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
Published June 2, updated August 3, 2026; tasks include a 1,200-line Python refactor, screenshot plus stack-trace debugging, and a 1,000-step Postgres migration.
Published conditions
Each model runs five times at temperature 0.2; the long-horizon task is unattended for four hours and quality is scored 1–5 by a senior reviewer.

Key data and applicable tasks

One-sentence takeaway

Using the same prompt, medians from five runs, and a senior reviewer, Ofox compared Max and Plus: Max had small quality/speed advantages on pure text and long-horizon migration, while Plus cost about five times less across the three tasks and supported visual input.

Test environment

  • Interface: Ofox OpenAI-compatible API.

  • Configuration: Same prompt, temperature 0.2; each task was run 5 times and the median was reported; quality was scored 1–5 by a senior reviewer.

  • Tasks: Asynchronous refactoring of a 1,200-line Python service; debugging a flaky test from a screenshot + stack trace; and a 1,000-step autonomous CLI migration from Postgres 14 to 16.

  • Long-horizon task: Each model ran unattended for 4 hours, below the 35-hour ceiling mentioned on the page.

Inputs/configuration

The article publishes task descriptions, input/output tokens, duration, tool calls, error recovery, and cost, but not the complete code repository, per-round tool traces, random seeds, reviewer rubric details, or raw API responses. Its method and budget order of magnitude can be replicated, but it cannot be recalculated line by line.

Results data

Task 1: Asynchronous refactoring of a 1,200-line Python service

MetricQwen3.7-PlusQwen3.7-Max
Input tokens12,84012,840
Output tokens4,2103,980
Median time47s41s
Quality4/54/5
Diff applied cleanlyYesYes
Estimated task cost$0.012$0.062

Task 2: Debugging a flaky test from a screenshot + stack trace

MetricQwen3.7-PlusQwen3.7-Max
Input8,420 + 1 image8,420, image dropped
Output tokens1,8302,140
Time12s9s
Quality5/52/5
Found the actual causeYesNo

Task 3: 1,000-step Postgres migration Agent

MetricQwen3.7-PlusQwen3.7-Max
Tool calls342351
Errors recovered4/55/5
Completion96%100%
Total cost$0.34$1.71

Conclusions

  • In the pure-text async refactor, both models scored 4/5; Max's median time was about 14% faster, but the task cost was about 5 times that of Plus.

  • The visual debugging task is not “cheaper Max vs. more expensive Plus”: Max could not accept the screenshot, while Plus found the cause at 5/5 and Max scored 2/5 and guessed the wrong location.

  • In the long-horizon migration, Max completed 100% with 5/5 error recovery, versus Plus at 96% and 4/5; for irreversible production migrations, a small quality difference may be worth paying for, while Plus has a cost advantage in rollback-capable staging.

  • The same article's pricing table lists Max input/output at $2.50/$7.50 and Plus at $0.40/$1.60; actual quotes should follow the current Alibaba Cloud page.

Limitations

  • Ofox provides both comparison APIs and promotional services, creating a platform interest; the task set has only 3 tasks and cannot represent every coding/Agent workflow.

  • The senior reviewer, sample code, and raw traces are not public, so quality scores may contain subjective elements.

  • Plus's visual advantage depends on image input through the API and the task design; Max's text path is not strictly equivalent to the visual path.

  • A 4-hour run and 342–351 tool calls do not prove a 35-hour ceiling or stability for arbitrary Agents.

Reproduction steps

  1. Use the same dated snapshot, prompt, temperature 0.2, and tool permissions to run Max and Plus separately through your own API/gateway.

  2. Repeat each task at least 5 times, saving input/output tokens, TTFT, wall-clock time, tool calls, error recovery, and the final diff.

  3. Have two reviewers who do not know the model identity score the results with a fixed rubric, and independently record whether the actual root cause was found and whether the tests passed.

  4. Separate image-input tasks from pure-text tasks; do not misclassify “the model cannot see the image” as weak reasoning ability.

  5. Calculate the cost of each result that passes without manual rework, then decide whether Max's quality premium is justified.

Source excerpt or observation (brief excerpt for compliance only)

The article's cost conclusion is “Plus wins on text-only cost,” but the 100% vs. 96% long-horizon migration result is a reminder to route irreversible tasks according to the cost of failure.

What this supports

  • It supports a controlled three-task cost and quality comparison under Ofox’s stated setup.

What this does not support

  • It cannot represent all repositories, longer autonomous runs, or current provider pricing outside those snapshots.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Ofox AI · Ofox (the article page does not show an individual author) · Original publication date 2026-06-02 · Site edit date 2026-09-20

Open original source

Qwen3.7 Max

Compare Qwen3.7 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.7 Max: What It Is, Access Routes, and Where It Fits

A sourced Qwen3.7 Max overview covering the dated snapshot, one-million-token API boundary, agentic use cases, pricing separation, and practical risks.

Related reviews

Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization ExperimentQwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed LedgerBenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.Qwen3.7-Max: Artificial Analysis Intelligence Index, Cost, and Speed BenchmarkArtificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.Qwen3.7-Max ITBench-AA Enterprise IT Operations and SRE Root-Cause Analysis BenchmarkITBench-AA evaluates Qwen3.7 Max on offline incident snapshots containing Prometheus, OpenTelemetry, Kubernetes, logs, and topology evidence in the Stirrup sandbox; tools and payload shape the result.Qwen3.7-Max: Long-Horizon Agents, Frontend Prototypes, and Office PromptsQwen’s long-horizon case separates frontend, office, and multi-step agent work into verifiable stages, with tools, timeouts, and final acceptance recorded separately.Qwen3.7-Max: Alibaba Cloud Model Studio Versions, Pricing, and Cache ConfigurationThe Model Studio page separates Qwen3.7 Max snapshots, the million-token context, cache billing, and regional limits so a test can fix version and cost assumptions first.Qwen3.7-Max: Three.js Electronic Rubik's Cube and 3D Physics Interaction Prototype PromptAlibaba Cloud’s public Three.js case turns visual references, interaction rules, and physics checks into an electronic cube prototype task for browser-based visual prototyping.Qwen3.7-Max: Long-Horizon Agent Prompts and Acceptance Closed Loop for GPU Kernel OptimizationThe Qwen team’s GPU-kernel case connects performance hypotheses, compilation tests, benchmark regressions, and long-horizon progress logs; hardware and verifier scope are critical boundaries.