Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elapsed time per task all increased substantially. The full-pass rate rose from 25.5% to 47.6%, which captures the change in long-chain tasks better than average accuracy alone.
Legal Research Bench: 208 isolated questions; accuracy 78.4% → 86.1%.
Full-pass rate: 25.5% → 47.6%.
Turns per task: 19.7 → 35.5; tool calls: 38 → 60; sources: 7.1 → 12.2; elapsed time: 809 seconds → 3,678 seconds.
CourtListener calls: 6.8 → 22.9; general web-search usage declined.
Maximum output tokens increased from 64K to 128K; legal research uses about 6× more reasoning, and final answers are 1.7× longer.
Task cost: Qwen3.8-Max $2.49; Opus 5 $6.76; Fable 5 $9.79; GPT-5.6 Sol $21.61.
It is more accurate to describe Qwen3.8-Max's advantage as an improvement in “long-chain completion rate/evidence coverage” than simply as being “smarter”; latency, tool calls, and cost must be stated alongside it. For scenarios that require fast responses, the highest reasoning tier should not be assumed by default.
Qwen 3.8 Max nearly doubled its score on Legal Research Bench in under three months, climbing from #22 to #4. This open-… This is a necessary excerpt; read the original source for full context.
The English above is the thread text corresponding to “Show original” on the X page; the original post also includes engagement data and comments, but navigation, ads, and unrelated recommendations were not included in the body.
Qwen3.8 Max