Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.7 Max · Media / benchmark · Independent measurement

Qwen3.7-Max ITBench-AA Enterprise IT Operations and SRE Root-Cause Analysis Benchmark

ITBench-AA evaluates Qwen3.7 Max on offline incident snapshots containing Prometheus, OpenTelemetry, Kubernetes, logs, and topology evidence in the Stirrup sandbox; tools and payload shape the result.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
The May 28, 2026 IBM/Artificial Analysis evaluation uses offline incident snapshots with metrics, traces, Kubernetes events, logs, and service topology.
Published conditions
The model gets shell and file-retrieval access in Stirrup; the result is not a tool-free chat score and full traces and deployment cost are not public.

Key data and applicable tasks

One-sentence takeaway

In ITBench-AA — the real-world enterprise-grade SRE operations benchmark jointly launched by IBM Research and Artificial Analysis — Qwen3.7-Max debuted at #3 globally upon release, demonstrating outstanding cross-system root-cause localization capabilities during multidimensional snapshot analysis of complex sandboxed Kubernetes incidents.

Test environment

  • Evaluation benchmark: ITBench-AA (independent reproduction edition of IBM ITBench) .

  • Test scenarios: 59 real-world enterprise-grade Kubernetes troubleshooting tasks (40 public scenarios + 19 unreleased private validation scenarios) , with each task tested 3 times repeatedly to eliminate random error.

  • Input data payload: Offline incident snapshots (including full Prometheus metrics, OpenTelemetry Traces, K8s Events, system alert logs, and application service topology graphs) .

  • Execution framework: Stirrup (open-source Agent execution sandbox, granting models autonomous Shell interaction and file retrieval permissions) .

  • Scoring criterion: Average Precision at Full Recall (calculating TP / (TP + FP) under the premise of zero false negatives) .

Inputs/configuration

  • Diagnostic output: The Agent is required to complete trace troubleshooting inside the sandbox and output a standardized JSON incident diagnostic report explicitly specifying root-cause entities (Deployment, Service, Pod, Namespace, NetworkPolicy, ConfigMap, etc.) .

  • Thinking mode: Thinking mode enabled, granting tool exploration and log filtering capabilities.

Results data

1. ITBench-AA Core Rankings and Accuracy Comparison (SRE Tasks)

ModelReasoning tierITBench-AA Composite PrecisionAverage Interaction Turns (Turns)Single-Task Inference Latency (Min)
GPT-5.6 Solmax56.2%30.85.18
GPT-5.6 Terramax51.0%37.73.74
Qwen3.7-Maxreasoning**Debut Top 3 (#3) **~35-40~3.5-4.0
Kimi K3max47.7%39.013.37
Claude Opus 4.7max46.7%68.210.16
GPT-5.5xhigh45.8%30.93.55
GPT-5.6 Lunamax40.3%45.93.77
Claude 4.5 Haikureasoning27.3%40.62.65
gpt-oss-120bhigh5.6%74.33.10
Nemotron 3 Superdefault1.1%98.85.61

2. Classic Incident Diagnosis Case Studies

  • Case study: Feature Flag configuration triggers downstream service avalanche:

    • Failure symptom: Sudden CPU spike and latency breach in the advertising service under the otel-demo namespace.

    • Surface-level trigger: Downstream Ad Deployment Pod triggered high-load alerts.

    • Model performance: Qwen3.7-Max successfully traced upstream along the topology relationship, accurately pinpointing the abnormal flag in the flagd-config ConfigMap ( adHighCpu: true ) rather than merely flagging the affected downstream Pods, avoiding false attribution.

  • Case study: Environment variable port mismatch leads to communication breakdown:

    • Failure symptom: Shipping service failed to call Quote service.

    • Model performance: Through Trace analysis and Deployment environment variable comparison, the model accurately captured the configuration defect where QUOTE_ADDR was mistakenly written with an invalid port ( quote:0000 ) .

Conclusions

  • First-tier performance in enterprise IT and SRE scenarios: When facing hundreds of megabytes of multidimensional mixed inputs spanning metrics, logs, and trace topologies, Qwen3.7-Max effectively utilizes Shell tools for targeted grep and correlation analysis, with its overall troubleshooting efficiency and localization precision firmly placed among the global leaders (#3) .

  • Balance between interaction turns and latency: An average of 35–40 interaction turns and approximately 3.5 minutes of decoding latency makes it more concise and efficient than Claude Opus 4.7 (68.2 turns, 10.16 minutes) .

Limitations

  • Even for the top three models, overall precision on ITBench-AA has not yet broken through 60%, indicating that fully autonomous enterprise-grade SRE operations remain in a high-difficulty frontier stage; AI is currently suited only as an auxiliary troubleshooting Copilot and should not be directly granted unreviewed self-healing write permissions on production clusters.

  • The evaluation is based on static offline incident snapshots and does not cover dynamic real-time fault injection or long-cycle closed-loop self-healing interactions.

Reproduction steps

  1. Clone the official evaluation suite github.com/ArtificialAnalysis/ITBench-AA and the execution engine Stirrup.

  2. Configure the sandbox environment to mount the Kubernetes incident snapshot data package provided by IBM.

  3. Point the endpoint to qwen3.7-max, run the complete 59-item SRE evaluation pipeline, and compare the generated JSON diagnostic reports against the Ground Truth.

Source excerpt or observation (brief excerpt for compliance only)

IBM Research and Artificial Analysis emphasize: “ITBench-AA tests AI agents on Kubernetes incident root-cause analysis... Qwen3.7-Max just hit #3 on ITbench-AA, demonstrating how well models handle real-world enterprise IT tasks, agentic-style”.

What this supports

  • ITBench-AA evaluates Qwen3.7 Max on offline incident snapshots containing Prometheus, OpenTelemetry, Kubernetes, logs, and topology evidence in the Stirrup sandbox; tools and payload shape the result.

What this does not support

  • ITBench-AA is bounded to its offline incident snapshots, Stirrup sandbox, tool permissions, and evaluator. It does not cover live production changes, organization-specific runbooks, or different telemetry quality.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis & IBM Research · Saurabh Jha et al. (IBM Research) & Artificial Analysis Team · Original publication date 2026-05-28 · Site edit date 2026-09-20

Open original source

Qwen3.7 Max

Compare Qwen3.7 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.7 Max: What It Is, Access Routes, and Where It Fits

A sourced Qwen3.7 Max overview covering the dated snapshot, one-million-token API boundary, agentic use cases, pricing separation, and practical risks.

Related reviews

Qwen3.7-Max: Official Complete Benchmarks and 35-Hour Autonomous Optimization ExperimentQwen reports Qwen3.7 Max results on coding, MCP/Skills, reasoning, multilingual, and long-horizon tool tasks, plus a 35-hour GPU optimization experiment; each score is harness-bound.Qwen3.7-Max: BenchLM Public Evidence Coverage and Speed LedgerBenchLM records 71.6/100, public rank #16/218, evidence-verified rank #13/104, 204 tok/s, and 14.07-second time to first token for Qwen3.7 Max; its aggregate definitions are mixed.Qwen3.7-Max vs. Qwen3.7-Plus: Cost and Quality on Three Real TasksOfox compares Qwen3.7 Max and Plus with the same prompts, five-run medians, and a senior reviewer; Max has small text and long-horizon advantages while Plus is about five times cheaper and accepts vision.Qwen3.7-Max: Artificial Analysis Intelligence Index, Cost, and Speed BenchmarkArtificial Analysis places Qwen3.7 Max in a 182-model pool and compares intelligence, cost per task, and output speed by reasoning setting; long reasoning traces change the cost materially.Qwen3.7-Max: Long-Horizon Agents, Frontend Prototypes, and Office PromptsQwen’s long-horizon case separates frontend, office, and multi-step agent work into verifiable stages, with tools, timeouts, and final acceptance recorded separately.Qwen3.7-Max: Alibaba Cloud Model Studio Versions, Pricing, and Cache ConfigurationThe Model Studio page separates Qwen3.7 Max snapshots, the million-token context, cache billing, and regional limits so a test can fix version and cost assumptions first.Qwen3.7-Max: Three.js Electronic Rubik's Cube and 3D Physics Interaction Prototype PromptAlibaba Cloud’s public Three.js case turns visual references, interaction rules, and physics checks into an electronic cube prototype task for browser-based visual prototyping.Qwen3.7-Max: Long-Horizon Agent Prompts and Acceptance Closed Loop for GPU Kernel OptimizationThe Qwen team’s GPU-kernel case connects performance hypotheses, compilation tests, benchmark regressions, and long-horizon progress logs; hardware and verifier scope are critical boundaries.