Qwen’s official release notes summarize multiple Qwen3.8 Max benchmarks; harnesses, samples, and reasoning parameters vary by task, so release and collection dates must remain separate rather than forming a current overall ranking.
Official Qwen Blog · Read evidenceQwen3.8 Max · Reviews and evidence
Which Qwen3.8 Max conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
Artificial Analysis separates Qwen3.8 Max quality, cost, speed, and verbosity; page version, reasoning tier, provider, and task sample need a fresh check, and the aggregate index must not become a cross-version trend.
Artificial Analysis · Read evidenceNYU Shanghai RITS material discusses Qwen3.8 Max agent turns and hallucination/cost proxies; task set, tools, repeats, and version follow the disclosed portion and cannot generalize to every agent workload.
NYU Shanghai RITS · Read evidenceFull reviews and related reading
Selected evidence
Qwen3.8 Max: Qwen3.8-Max: Official Release Notes and Complete Performance Results
Qwen’s official release notes summarize multiple Qwen3.8 Max benchmarks; harnesses, samples, and reasoning parameters vary by task, so release and collection dates must remain separate rather than forming a current overall ranking.
Unverified: the original source could not be rechecked.
- Source
- Official Qwen release notes; vendor-reported benchmarks
- Version
- Qwen3.8 Max exact snapshot follows the source
- Task sets
- Multiple benchmarks with non-uniform harnesses, samples, and reasoning settings
Qwen3.8 Max: Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and Verbosity
Artificial Analysis separates Qwen3.8 Max quality, cost, speed, and verbosity; page version, reasoning tier, provider, and task sample need a fresh check, and the aggregate index must not become a cross-version trend.
Unverified: the original source could not be rechecked.
- Platform/version
- Artificial Analysis index page; version and window need a fresh check
- Metrics
- Quality, cost, speed, and verbosity are interpreted separately
- Configuration
- Provider, tier, sample, and task set are incomplete
Qwen3.8 Max: Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination Cost
NYU Shanghai RITS material discusses Qwen3.8 Max agent turns and hallucination/cost proxies; task set, tools, repeats, and version follow the disclosed portion and cannot generalize to every agent workload.
Unverified: the original source could not be rechecked.
- Study
- NYU Shanghai RITS agentic index; version and harness follow the source
- Metrics
- Agent turns, hallucination, or cost proxies, not one quality score
- Sample
- Task set, repeats, and tools follow the disclosed portion
Qwen3.8 Max: Qwen3.8 Max: BenchLM's Source-Verifiable Benchmark Ledger
BenchLM separates Qwen3.8 Max exact-source benchmark rows from its aggregate ranking; weights, providers, harnesses, samples, and dates differ, making it a verifiable ledger rather than a unified independent rerun.
Unverified: the original source could not be rechecked.
- Platform
- BenchLM exact-source rows are separate from aggregate ranking
- Coverage
- Category weights and directory are dynamic
- Configuration
- Provider, harness, samples, and dates differ by benchmark
All sources
All sources
Qwen3.8 Max: Qwen3.8-Max: Official Release Notes and Complete Performance Results
Qwen’s official release notes summarize multiple Qwen3.8 Max benchmarks; harnesses, samples, and reasoning parameters vary by task, so release and collection dates must remain separate rather than forming a current overall ranking.
Unverified: the original source could not be rechecked.
- Source
- Official Qwen release notes; vendor-reported benchmarks
- Version
- Qwen3.8 Max exact snapshot follows the source
- Task sets
- Multiple benchmarks with non-uniform harnesses, samples, and reasoning settings
Qwen3.8 Max: Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and Verbosity
Artificial Analysis separates Qwen3.8 Max quality, cost, speed, and verbosity; page version, reasoning tier, provider, and task sample need a fresh check, and the aggregate index must not become a cross-version trend.
Unverified: the original source could not be rechecked.
- Platform/version
- Artificial Analysis index page; version and window need a fresh check
- Metrics
- Quality, cost, speed, and verbosity are interpreted separately
- Configuration
- Provider, tier, sample, and task set are incomplete
Qwen3.8 Max: Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination Cost
NYU Shanghai RITS material discusses Qwen3.8 Max agent turns and hallucination/cost proxies; task set, tools, repeats, and version follow the disclosed portion and cannot generalize to every agent workload.
Unverified: the original source could not be rechecked.
- Study
- NYU Shanghai RITS agentic index; version and harness follow the source
- Metrics
- Agent turns, hallucination, or cost proxies, not one quality score
- Sample
- Task set, repeats, and tools follow the disclosed portion
Qwen3.8 Max: Qwen3.8 Max: BenchLM's Source-Verifiable Benchmark Ledger
BenchLM separates Qwen3.8 Max exact-source benchmark rows from its aggregate ranking; weights, providers, harnesses, samples, and dates differ, making it a verifiable ledger rather than a unified independent rerun.
Unverified: the original source could not be rechecked.
- Platform
- BenchLM exact-source rows are separate from aggregate ranking
- Coverage
- Category weights and directory are dynamic
- Configuration
- Provider, harness, samples, and dates differ by benchmark
Qwen3.8 Max: Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research Bench
Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number of turns, tool calls, sources, and elap。
Unverified: the original source could not be rechecked.
- Model/version
- Qwen3.8-Max; source title “Qwen3.8 Max: Qwen3.8-Max: Persistence, Full-pass Rate, and Task Cost on Legal Research Bench”. Exact snapshot follows the original source.
- Task/harness
- Vals AI's thread explains Qwen3.8-Max's improvement as being “more persistent,” not simply “smarter”: its Legal Research Bench ranking rose from No. 22 to No. 4, at an approximate cost of $2.49 per task, while the number The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Qwen3.8 Max: Qwen3.8-27B Local Quantized Model: Reasoning Effort Level Test
The author tested Qwen3.8-27B on four machines: MLX 4-bit on an M5 Max, and unsloth/Qwen3.8-27B-NVFP4 running through vLLM on a DGX Spark. He observed a marked jump from thinking off to effort=low, but on the 4-bit model, xhigh can take an extreme amount of ti。
Unverified: the original source could not be rechecked.
- Model/version
- Qwen3.8-Max; source title “Qwen3.8 Max: Qwen3.8-27B Local Quantized Model: Reasoning Effort Level Test”. Exact snapshot follows the original source.
- Task/harness
- The author tested Qwen3.8-27B on four machines: MLX 4-bit on an M5 Max, and unsloth/Qwen3.8-27B-NVFP4 running through vLLM on a DGX Spark. He observed a marked jump from thinking off to effort=low, but on the 4-bit model The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Qwen3.8 Max: Reddit Community: Qwen3.8-Max Coding Ability, Speed, and Usage Quota
This is a community discussion asking whether Qwen3.8-Max is really suitable for programming. The feedback is polarized: some users consider it close to Claude/GPT, while others find it slow, expensive, and prone to overthinking. Another user used it to genera。
Unverified: the original source could not be rechecked.
- Model/version
- Qwen3.8-Max; source title “Qwen3.8 Max: Reddit Community: Qwen3.8-Max Coding Ability, Speed, and Usage Quota”. Exact snapshot follows the original source.
- Task/harness
- This is a community discussion asking whether Qwen3.8-Max is really suitable for programming. The feedback is polarized: some users consider it close to Claude/GPT, while others find it slow, expensive, and prone to over The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Qwen3.8 Max: Reddit: Is Qwen3.8-Max's High Score Inflated by a Single Benchmark?
The original poster noticed that Qwen3.8-Max has a very high composite score, but did not feel equally intelligent while using it for research in the Qwen App, and asked whether “benchmaxxing” was involved. Replies pointed out that the composite score is pulle。
Unverified: the original source could not be rechecked.
- Model/version
- Qwen3.8-Max; source title “Qwen3.8 Max: Reddit: Is Qwen3.8-Max's High Score Inflated by a Single Benchmark?”. Exact snapshot follows the original source.
- Task/harness
- The original poster noticed that Qwen3.8-Max has a very high composite score, but did not feel equally intelligent while using it for research in the Qwen App, and asked whether “benchmaxxing” was involved. Replies point The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Qwen3.8 Max: Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test
One-sentence takeaway On the same 269-file, 60-minute StackPerf codebase architecture task using OpenCode 1.17.13, Qwen3.8-Max Preview scored 80 and Kimi K3 scored 83; Qwen was stronger on system boundaries, evidence citations, and replay metadata, while Kimi 。
Unverified: the original source could not be rechecked.
- Model/version
- Qwen3.8-Max; source title “Qwen3.8 Max: Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test”. Exact snapshot follows the original source.
- Task/harness
- One-sentence takeaway On the same 269-file, 60-minute StackPerf codebase architecture task using OpenCode 1.17.13, Qwen3.8-Max Preview scored 80 and Kimi K3 scored 83; Qwen was stronger on system boundaries, evidence cit The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Qwen3.8 Max
Compare Qwen3.8 Max in Tabbit
Model access, features, and permissions depend on your current client account.