Qwen3.8 Max · Community source · Personal experience
The original poster noticed that Qwen3.8-Max has a very high composite score, but did not feel equally intelligent while using it for research in the Qwen App, and asked whether “benchmaxxing” was involved. Replies pointed out that the composite score is pulle。
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
The original poster noticed that Qwen3.8-Max has a very high composite score, but did not feel equally intelligent while using it for research in the Qwen App, and asked whether “benchmaxxing” was involved. Replies pointed out that the composite score is pulled upward by agentic tool-use projects such as Tau3-banking; others argued that most new models can handle common coding tasks when given enough context and clear prompts, and that the real-world difference between 50 and 60 points cannot be inferred from the score alone.
Examine component metrics and task definitions first; do not cite only the composite score.
The weighting of tool-use benchmarks can materially change the total; record the tools, prompts, context, and pass criteria.
Community day-to-day feedback suggests it is suitable for general coding and agent tasks, but it cannot replace regression testing on real projects.
Can someone tell me is it really this good? Because when I try it on qwen app it doesn't even enough smart for some… This is a necessary excerpt; read the original source for full context.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
reddit.com · Author not disclosed · Original publication date Unknown · Site edit date 2026-09-20
Open original sourceQwen3.8 Max
Download the Tabbit client to check model access