Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Qwen3.8 Max · Media / benchmark · Independent measurement

Qwen3.8 Max: Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test

One-sentence takeaway On the same 269-file, 60-minute StackPerf codebase architecture task using OpenCode 1.17.13, Qwen3.8-Max Preview scored 80 and Kimi K3 scored 83; Qwen was stronger on system boundaries, evidence citations, and replay metadata, while Kimi 。

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Model/version
Qwen3.8-Max; source title “Qwen3.8 Max: Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test”. Exact snapshot follows the original source.
Task/harness
One-sentence takeaway On the same 269-file, 60-minute StackPerf codebase architecture task using OpenCode 1.17.13, Qwen3.8-Max Preview scored 80 and Kimi K3 scored 83; Qwen was stronger on system boundaries, evidence cit The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.

Key data and applicable tasks

One-sentence takeaway

On the same 269-file, 60-minute StackPerf codebase architecture task using OpenCode 1.17.13, Qwen3.8-Max Preview scored 80 and Kimi K3 scored 83; Qwen was stronger on system boundaries, evidence citations, and replay metadata, while Kimi was stronger on revision, regeneration, and scenario lifecycle, and both still required independent factual verification.

Use cases

  • Suitable tasks: Long-context architecture reviews that require reading an unfamiliar codebase, comparing system responsibilities, designing data contracts, writing migration plans, and maintaining an evidence ledger.

  • Unsuitable tasks: Treating one codebase architecture blind test as a general coding benchmark, single-turn Q&A, or the current score for Qwen3.8-Max GA.

  • Applicable model versions: Qwen3.8-Max-Preview (the 2026-07-19 service snapshot), not a direct substitute for later GA versions.

  • Applicable clients, Agents, or APIs: OpenCode 1.17.13; Qwen used Alibaba's international Token Plan endpoint, and Kimi used the Kimi Code subscription endpoint.

  • Recommended reasoning tier and parameters: Qwen enable_thinking: true, with no fixed thinking budget; the article does not disclose temperature or the complete system prompt, which should not be filled in speculatively.

Test environment and input/configuration

  • The input consisted of read-only snapshots of two unfamiliar projects, trilogy-group/ttv-pipeline and kumanday/media-tooling, totaling 269 files; generated media, caches, virtual environments, and internal repository directories were excluded.

  • The task required file-by-file/line-by-line evidence analysis, distinguishing global planning from provider-specific generation, comparing 15/30-second video generation, designing typed/JSON contracts, and proposing migration phases, tests, risks, and an evidence ledger.

  • Both models used the same task card, permissions, and 60-minute wall-clock limit, with a completion ceiling of 65,536 tokens.

  • Read, search, listing, and safety checks were allowed; editing, network access, writing source files, plugins, and subagents were prohibited.

  • StackPerf registered the sessions and collected request, tool, and output metrics; before blind evaluation, the reports were renamed Report A/B, followed by a factual verification stage.

Results data

  • Qwen3.8-Max Preview: 80/100 after factual deductions; Kimi K3: 83/100.

  • Qwen: 22 gateway requests and 44 tool calls, all tool calls successful; the report was longer, with 354 citations, and verification found no fabricated paths or symbols, but 2 cited line numbers were inaccurate and 7 groups of conclusions exceeded the code evidence.

  • Kimi: 53 tool calls, with 2 compound shell commands rejected by policy before recovery; it had 274 citations, no fabricated paths or symbols were found, and it likewise had 7 groups of conclusions that exceeded the evidence.

  • Both models reached the same system boundary: one side owned global planning and final assembly, while the other handled provider-specific generation, retries, and provenance, connected through a versioned contract.

  • Qwen's GenerationManifest / GenerationResult and replay records were more complete, recording model, seed, prompt, references, provider request ID, media URI, duration, and cost; Kimi modeled revision, supersession, retry history, and take invalidation more completely.

  • The cache hit rate for repeated prompts on both provider paths exceeded 90%; this figure is affected by the provider, routing, and session cache.

Workflow/reproduction steps

  1. Fix SHA-256 snapshots of the two projects, excluding media, caches, and internal repository directories.

  2. Load the same task card in OpenCode 1.17.13, set the same permissions, 60-minute limit, and 65,536-token completion ceiling, and retain each model's native reasoning settings.

  3. Prohibit editing, network access, plugins, and subagents; allow only safe reading, retrieval, and directory checks.

  4. Use StackPerf to record every request, tool call, failure, token count, and cache event; anonymize the two final reports for blind evaluation, then perform factual verification of paths, symbols, and conclusions.

  5. Report scores, call efficiency, citation coverage, and unsupported claims separately; do not extrapolate a single 80/83 result into an overall model ranking.

Conclusions

This is a joint measurement of “model × harness × task,” not proof of Qwen3.8-Max's absolute capability. Qwen had advantages in architecture boundaries, replay, and fewer tool calls; Kimi had advantages in lifecycle and regeneration design. For high-value code review, having two models produce outputs in parallel and then cross-checking them is safer than sending one model's polished long report directly into implementation.

Limitations

  • The test used Qwen3.8-Max Preview; the article explicitly says to rerun it after changes to the GA version, official API, or open weights.

  • Each model had only one complete service path, one task, and one 60-minute session, so the effects of model capability cannot be separated from the endpoint, cache, rate limit, and harness.

  • The blind-evaluation score included human factual deductions, and the article does not disclose the complete task prompt, all tool schemas, or the raw turn-by-turn transcript.

  • The task concerned video-system architecture and codebase reading; it cannot replace SWE-bench, Terminal-Bench, Chinese-language, or multimodal regression tests.

Original evidence and data

The article makes public links to the two projects and the StackPerf repository, the 269-file scale, OpenCode version, permissions, time/token limits, model reasoning configuration, request/tool counts, citation counts, and factual-verification results; these fields are sufficient to reproduce the experiment's structure, but not the exact score of 80 without the same snapshots and service access.

Source excerpt or observation (short quote for compliance only)

The article describes Qwen's configuration as “enable_thinking: true and no fixed thinking budget”; this demonstrates the test configuration only and does not represent the default settings of every Qwen3.8-Max client.

What this supports

  • Supports the source-specific observation in “Qwen3.8 Max: Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test”: One-sentence takeaway On the same 269-file, 60-minute StackPerf codebase architecture task using OpenCode 1.17.13, Qwen3.8-Max Preview scored 80 and Kimi K3 scored 83; Qwen was stronger on s

What this does not support

  • Does not support a general capability, production success-rate, or current-ranking claim from “Qwen3.8 Max: Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test”; the source lacks a controlled task set, provider snapshot, and repeated independent retest.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Trilogy AI Center of Excellence (Substack) · Leonardo Gonzalez · Original publication date 2026-07-19 · Site edit date 2026-09-20

Open original source

Qwen3.8 Max

Compare Qwen3.8 Max in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Qwen3.8 Max: What Changed, What It Costs, and Who It Fits

A sourced Qwen3.8 Max overview covering the 0902 snapshot, multimodal boundary, benchmark caveats, access routes and a safer pilot.

Related reviews

Qwen3.8 Max: Reddit Community: Qwen3.8-Max Coding Ability, Speed, and Usage QuotaThis is a community discussion asking whether Qwen3.8-Max is really suitable for programming. The feedback is polarized: some users consider it close to Claude/GPT, while others find it slow, expensive, and prone to overthinking. Another user used it to genera。Qwen3.8 Max: Qwen3.8-Max: Official Release Notes and Complete Performance ResultsQwen’s official release notes summarize multiple Qwen3.8 Max benchmarks; harnesses, samples, and reasoning parameters vary by task, so release and collection dates must remain separate rather than forming a current overall ranking.Qwen3.8 Max: Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and VerbosityArtificial Analysis separates Qwen3.8 Max quality, cost, speed, and verbosity; page version, reasoning tier, provider, and task sample need a fresh check, and the aggregate index must not become a cross-version trend.Qwen3.8 Max: Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination CostNYU Shanghai RITS material discusses Qwen3.8 Max agent turns and hallucination/cost proxies; task set, tools, repeats, and version follow the disclosed portion and cannot generalize to every agent workload.Qwen3.8 Max: Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting GuideTurn Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling GuideTurn Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission BoundariesTurn Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission Boundaries into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Qwen3.8 Max: Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” ReportsTurn Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.