Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityQwen3.8 Max

Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research Bench

Original source

x.com

AuthorVals AI

Source date2026-08-12

Tabbit curation2026-08-19

Read original

Summary

Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elapsed time per task all increased substantially. The full-pass rate rose from 25.5% to 47.6%, which captures the change in long-chain tasks better than average accuracy alone.

Key data

  • Legal Research Bench: 208 isolated questions; accuracy 78.4% → 86.1%.

  • Full-pass rate: 25.5% → 47.6%.

  • Turns per task: 19.7 → 35.5; tool calls: 38 → 60; sources: 7.1 → 12.2; elapsed time: 809 seconds → 3,678 seconds.

  • CourtListener calls: 6.8 → 22.9; general web-search usage declined.

  • Maximum output tokens increased from 64K to 128K; legal research uses about 6× more reasoning, and final answers are 1.7× longer.

  • Task cost: Qwen3.8-Max $2.49; Opus 5 $6.76; Fable 5 $9.79; GPT-5.6 Sol $21.61.

Implications for Tabbit international content

It is more accurate to describe Qwen3.8-Max's advantage as an improvement in “long-chain completion rate/evidence coverage” than simply as being “smarter”; latency, tool calls, and cost must be stated alongside it. For scenarios that require fast responses, the highest reasoning tier should not be assumed by default.

Article text

Qwen 3.8 Max nearly doubled its score on Legal Research Bench in under three months, climbing from #22 to #4. This open-… This is a necessary excerpt; read the original source for full context.

Collection notes

The English above is the thread text corresponding to “Show original” on the X page; the original post also includes engagement data and comments, but navigation, ads, and unrelated recommendations were not included in the body.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Qwen3.8 Max

Use and compare models in Tabbit

Qwen3.8 Max

Related reviews

MediaOfficial Qwen Blog2026-08-03

Qwen3.8-Max: Official Release Notes and Complete Performance Results

MediaArtificial Analysis

Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and Verbosity

MediaNYU Shanghai RITS

Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination Cost

MediaTrilogy AI Center of Excellence (Substack)2026-07-19

Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test

Qwen3.8 Max

Related prompts

Mediaqwen.ai2026-08-03

Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide

MediaEvoLink.AI Blog2026-08-03

Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide

CommunityReddit, r/QwenAI2026-08-09

Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission Boundaries

CommunityReddit, r/QwenAI2026-08-06

Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports