Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.8 Max · Community source · Independent measurement

Qwen3.8 Max: Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research Bench

Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elap。

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceIndependent measurementEdited 2026-09-20

Test conditions

Model/version
Qwen3.8-Max; source title “Qwen3.8 Max: Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research Bench”. Exact snapshot follows the original source.
Task/harness
Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.

Key data and applicable tasks

Summary

Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elapsed time per task all increased substantially. The full-pass rate rose from 25.5% to 47.6%, which captures the change in long-chain tasks better than average accuracy alone.

Key data

  • Legal Research Bench: 208 isolated questions; accuracy 78.4% → 86.1%.

  • Full-pass rate: 25.5% → 47.6%.

  • Turns per task: 19.7 → 35.5; tool calls: 38 → 60; sources: 7.1 → 12.2; elapsed time: 809 seconds → 3,678 seconds.

  • CourtListener calls: 6.8 → 22.9; general web-search usage declined.

  • Maximum output tokens increased from 64K to 128K; legal research uses about 6× more reasoning, and final answers are 1.7× longer.

  • Task cost: Qwen3.8-Max $2.49; Opus 5 $6.76; Fable 5 $9.79; GPT-5.6 Sol $21.61.

Implications for Tabbit international content

It is more accurate to describe Qwen3.8-Max's advantage as an improvement in “long-chain completion rate/evidence coverage” than simply as being “smarter”; latency, tool calls, and cost must be stated alongside it. For scenarios that require fast responses, the highest reasoning tier should not be assumed by default.

Article text

Qwen 3.8 Max nearly doubled its score on Legal Research Bench in under three months, climbing from #22 to #4. This open-… This is a necessary excerpt; read the original source for full context.

Collection notes

The English above is the thread text corresponding to “Show original” on the X page; the original post also includes engagement data and comments, but navigation, ads, and unrelated recommendations were not included in the body.

What this supports

  • Supports the source-specific observation in “Qwen3.8 Max: Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research Bench”: Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2

What this does not support

  • Does not support a general capability, production success-rate, or current-ranking claim from “Qwen3.8 Max: Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research Bench”; the source lacks a controlled task set, provider snapshot, and repeated independent retest.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

x.com · Vals AI · Original publication date 2026-08-12 · Site edit date 2026-09-20

Open original source

Qwen3.8 Max

Compare Qwen3.8 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.8 Max: What Changed, What It Costs, and Who It Fits

A sourced Qwen3.8 Max overview covering the 0902 snapshot, multimodal boundary, benchmark caveats, access routes and a safer pilot.

Related reviews

Qwen3.8 Max: Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and VerbosityArtificial Analysis separates Qwen3.8 Max quality, cost, speed, and verbosity; page version, reasoning tier, provider, and task sample need a fresh check, and the aggregate index must not become a cross-version trend.Qwen3.8 Max: Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination CostNYU Shanghai RITS material discusses Qwen3.8 Max agent turns and hallucination/cost proxies; task set, tools, repeats, and version follow the disclosed portion and cannot generalize to every agent workload.Qwen3.8 Max: Qwen3.8 Max: BenchLM's Source-Verifiable Benchmark LedgerBenchLM separates Qwen3.8 Max exact-source benchmark rows from its aggregate ranking; weights, providers, harnesses, samples, and dates differ, making it a verifiable ledger rather than a unified independent rerun.Qwen3.8 Max: Reddit Community: Qwen3.8-Max Coding Ability, Speed, and Usage QuotaThis is a community discussion asking whether Qwen3.8-Max is really suitable for programming. The feedback is polarized: some users consider it close to Claude/GPT, while others find it slow, expensive, and prone to overthinking. Another user used it to genera。Qwen3.8 Max: Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” ReportsTurn Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting GuideTurn Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling GuideTurn Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Roleplay: Direction Following, Reasoning Time, and Preset FeedbackTurn Qwen3.8-Max Roleplay: Direction Following, Reasoning Time, and Preset Feedback into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.