Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.8 Max · Community source · Personal experience

Qwen3.8 Max: Qwen3.8-27B Local Quantized Model: Reasoning Effort Level Test

The author tested Qwen3.8-27B on four machines: MLX 4-bit on an M5 Max, and unsloth/Qwen3.8-27B-NVFP4 running through vLLM on a DGX Spark. He observed a marked jump from thinking off to effort=low, but on the 4-bit model, xhigh can take an extreme amount of ti。

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model/version
Qwen3.8-Max; source title “Qwen3.8 Max: Qwen3.8-27B Local Quantized Model: Reasoning Effort Level Test”. Exact snapshot follows the original source.
Task/harness
The author tested Qwen3.8-27B on four machines: MLX 4-bit on an M5 Max, and unsloth/Qwen3.8-27B-NVFP4 running through vLLM on a DGX Spark. He observed a marked jump from thinking off to effort=low, but on the 4-bit model The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.

Key data and applicable tasks

Summary

The author tested Qwen3.8-27B on four machines: MLX 4-bit on an M5 Max, and unsloth/Qwen3.8-27B-NVFP4 running through vLLM on a DGX Spark. He observed a marked jump from thinking off to effort=low, but on the 4-bit model, xhigh can take an extreme amount of time or be truncated because of long reasoning and quantization error. Follow-up replies also provided information about a 64K budget, an 8-bit retest, and memory consumption.

Key facts

  • Results were averaged across two runs, with each run lasting several hours.

  • effort=low trades higher token costs for results approaching frontier-model quality.

  • On the 4-bit model, xhigh may exceed 50K thinking tokens without completing; the author used geometric means to exclude extreme values.

  • The author later said that a 64K budget allowed some xhigh tasks to complete; some tests took about an hour but still reached 100%.

  • The M5 Max has 128GB of RAM; the 4-bit tests used about 23–24GB, while 8-bit used about 40GB; longer contexts require additional memory.

Assessment

This is not a direct evaluation of the Qwen3.8-Max cloud model, but a local quantized test of an open-weights 27B model from the same family. Its implications for prompting and runtime configuration are to establish cost and latency baselines with low/medium first, then raise effort separately for difficult tasks, with explicit token and time limits for xhigh.

Article text

Qwen3.8-27B - Interesting test results on effort I ran tests on 4 machines overnight and have been tweaking the tests al… This is a necessary excerpt; read the original source for full context.

Thread additions

I ran the model with a 64k token budget and eventually had to use that for xhigh to complete the tasks that were failing… This is a necessary excerpt; read the original source for full context.

What this supports

  • Supports the source-specific observation in “Qwen3.8 Max: Qwen3.8-27B Local Quantized Model: Reasoning Effort Level Test”: The author tested Qwen3.8-27B on four machines: MLX 4-bit on an M5 Max, and unsloth/Qwen3.8-27B-NVFP4 running through vLLM on a DGX Spark. He observed a marked jump from thinking off to effo

What this does not support

  • Does not support a general capability, production success-rate, or current-ranking claim from “Qwen3.8 Max: Qwen3.8-27B Local Quantized Model: Reasoning Effort Level Test”; the source lacks a controlled task set, provider snapshot, and repeated independent retest.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

x.com · John T Davies · Original publication date 2026-08-16 · Site edit date 2026-09-20

Open original source

Qwen3.8 Max

Compare Qwen3.8 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.8 Max: What Changed, What It Costs, and Who It Fits

A sourced Qwen3.8 Max overview covering the 0902 snapshot, multimodal boundary, benchmark caveats, access routes and a safer pilot.

Related reviews

Qwen3.8 Max: Qwen3.8-Max: Official Release Notes and Complete Performance ResultsQwen’s official release notes summarize multiple Qwen3.8 Max benchmarks; harnesses, samples, and reasoning parameters vary by task, so release and collection dates must remain separate rather than forming a current overall ranking.Qwen3.8 Max: Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and VerbosityArtificial Analysis separates Qwen3.8 Max quality, cost, speed, and verbosity; page version, reasoning tier, provider, and task sample need a fresh check, and the aggregate index must not become a cross-version trend.Qwen3.8 Max: Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination CostNYU Shanghai RITS material discusses Qwen3.8 Max agent turns and hallucination/cost proxies; task set, tools, repeats, and version follow the disclosed portion and cannot generalize to every agent workload.Qwen3.8 Max: Qwen3.8 Max: BenchLM's Source-Verifiable Benchmark LedgerBenchLM separates Qwen3.8 Max exact-source benchmark rows from its aggregate ranking; weights, providers, harnesses, samples, and dates differ, making it a verifiable ledger rather than a unified independent rerun.Qwen3.8 Max: Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission BoundariesTurn Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission Boundaries into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” ReportsTurn Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting GuideTurn Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling GuideTurn Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.