Qwen3.8 Max · Community source · Independent measurement
Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elap。
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elapsed time per task all increased substantially. The full-pass rate rose from 25.5% to 47.6%, which captures the change in long-chain tasks better than average accuracy alone.
Legal Research Bench: 208 isolated questions; accuracy 78.4% → 86.1%.
Full-pass rate: 25.5% → 47.6%.
Turns per task: 19.7 → 35.5; tool calls: 38 → 60; sources: 7.1 → 12.2; elapsed time: 809 seconds → 3,678 seconds.
CourtListener calls: 6.8 → 22.9; general web-search usage declined.
Maximum output tokens increased from 64K to 128K; legal research uses about 6× more reasoning, and final answers are 1.7× longer.
Task cost: Qwen3.8-Max $2.49; Opus 5 $6.76; Fable 5 $9.79; GPT-5.6 Sol $21.61.
It is more accurate to describe Qwen3.8-Max's advantage as an improvement in “long-chain completion rate/evidence coverage” than simply as being “smarter”; latency, tool calls, and cost must be stated alongside it. For scenarios that require fast responses, the highest reasoning tier should not be assumed by default.
Qwen 3.8 Max nearly doubled its score on Legal Research Bench in under three months, climbing from #22 to #4. This open-… This is a necessary excerpt; read the original source for full context.
The English above is the thread text corresponding to “Show original” on the X page; the original post also includes engagement data and comments, but navigation, ads, and unrelated recommendations were not included in the body.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
x.com · Vals AI · Original publication date 2026-08-12 · Site edit date 2026-09-20
Open original sourceQwen3.8 Max
Download the Tabbit client to check model access