Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityQwen3.8 Max

Qwen3.8-27B Local Quantized Model: Reasoning Effort Level Test

Original source

x.com

AuthorJohn T Davies

Source date2026-08-16

Tabbit curation2026-08-19

Read original

Summary

The author tested Qwen3.8-27B on four machines: MLX 4-bit on an M5 Max, and unsloth/Qwen3.8-27B-NVFP4 running through vLLM on a DGX Spark. He observed a marked jump from thinking off to effort=low, but on the 4-bit model, xhigh can take an extreme amount of time or be truncated because of long reasoning and quantization error. Follow-up replies also provided information about a 64K budget, an 8-bit retest, and memory consumption.

Key facts

  • Results were averaged across two runs, with each run lasting several hours.

  • effort=low trades higher token costs for results approaching frontier-model quality.

  • On the 4-bit model, xhigh may exceed 50K thinking tokens without completing; the author used geometric means to exclude extreme values.

  • The author later said that a 64K budget allowed some xhigh tasks to complete; some tests took about an hour but still reached 100%.

  • The M5 Max has 128GB of RAM; the 4-bit tests used about 23–24GB, while 8-bit used about 40GB; longer contexts require additional memory.

Assessment

This is not a direct evaluation of the Qwen3.8-Max cloud model, but a local quantized test of an open-weights 27B model from the same family. Its implications for prompting and runtime configuration are to establish cost and latency baselines with low/medium first, then raise effort separately for difficult tasks, with explicit token and time limits for xhigh.

Article text

Qwen3.8-27B - Interesting test results on effort I ran tests on 4 machines overnight and have been tweaking the tests al… This is a necessary excerpt; read the original source for full context.

Thread additions

I ran the model with a 64k token budget and eventually had to use that for xhigh to complete the tasks that were failing… This is a necessary excerpt; read the original source for full context.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Qwen3.8 Max

Use and compare models in Tabbit

Qwen3.8 Max

Related reviews

MediaOfficial Qwen Blog2026-08-03

Qwen3.8-Max: Official Release Notes and Complete Performance Results

MediaArtificial Analysis

Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and Verbosity

MediaNYU Shanghai RITS

Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination Cost

MediaTrilogy AI Center of Excellence (Substack)2026-07-19

Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test

Qwen3.8 Max

Related prompts

Mediaqwen.ai2026-08-03

Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide

MediaEvoLink.AI Blog2026-08-03

Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide

CommunityReddit, r/QwenAI2026-08-09

Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission Boundaries

CommunityReddit, r/QwenAI2026-08-06

Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports