Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaQwen3.8 Max

Qwen3.8-Max Preview: Trilogy AI's StackPerf Codebase Architecture Blind Test

Original source

Trilogy AI Center of Excellence (Substack)

AuthorLeonardo Gonzalez

Source date2026-07-19

Tabbit curation2026-08-19

Read original

One-sentence takeaway

On the same 269-file, 60-minute StackPerf codebase architecture task using OpenCode 1.17.13, Qwen3.8-Max Preview scored 80 and Kimi K3 scored 83; Qwen was stronger on system boundaries, evidence citations, and replay metadata, while Kimi was stronger on revision, regeneration, and scenario lifecycle, and both still required independent factual verification.

Use cases

  • Suitable tasks: Long-context architecture reviews that require reading an unfamiliar codebase, comparing system responsibilities, designing data contracts, writing migration plans, and maintaining an evidence ledger.

  • Unsuitable tasks: Treating one codebase architecture blind test as a general coding benchmark, single-turn Q&A, or the current score for Qwen3.8-Max GA.

  • Applicable model versions: Qwen3.8-Max-Preview (the 2026-07-19 service snapshot), not a direct substitute for later GA versions.

  • Applicable clients, Agents, or APIs: OpenCode 1.17.13; Qwen used Alibaba's international Token Plan endpoint, and Kimi used the Kimi Code subscription endpoint.

  • Recommended reasoning tier and parameters: Qwen enable_thinking: true, with no fixed thinking budget; the article does not disclose temperature or the complete system prompt, which should not be filled in speculatively.

Test environment and input/configuration

  • The input consisted of read-only snapshots of two unfamiliar projects, trilogy-group/ttv-pipeline and kumanday/media-tooling, totaling 269 files; generated media, caches, virtual environments, and internal repository directories were excluded.

  • The task required file-by-file/line-by-line evidence analysis, distinguishing global planning from provider-specific generation, comparing 15/30-second video generation, designing typed/JSON contracts, and proposing migration phases, tests, risks, and an evidence ledger.

  • Both models used the same task card, permissions, and 60-minute wall-clock limit, with a completion ceiling of 65,536 tokens.

  • Read, search, listing, and safety checks were allowed; editing, network access, writing source files, plugins, and subagents were prohibited.

  • StackPerf registered the sessions and collected request, tool, and output metrics; before blind evaluation, the reports were renamed Report A/B, followed by a factual verification stage.

Results data

  • Qwen3.8-Max Preview: 80/100 after factual deductions; Kimi K3: 83/100.

  • Qwen: 22 gateway requests and 44 tool calls, all tool calls successful; the report was longer, with 354 citations, and verification found no fabricated paths or symbols, but 2 cited line numbers were inaccurate and 7 groups of conclusions exceeded the code evidence.

  • Kimi: 53 tool calls, with 2 compound shell commands rejected by policy before recovery; it had 274 citations, no fabricated paths or symbols were found, and it likewise had 7 groups of conclusions that exceeded the evidence.

  • Both models reached the same system boundary: one side owned global planning and final assembly, while the other handled provider-specific generation, retries, and provenance, connected through a versioned contract.

  • Qwen's GenerationManifest / GenerationResult and replay records were more complete, recording model, seed, prompt, references, provider request ID, media URI, duration, and cost; Kimi modeled revision, supersession, retry history, and take invalidation more completely.

  • The cache hit rate for repeated prompts on both provider paths exceeded 90%; this figure is affected by the provider, routing, and session cache.

Workflow/reproduction steps

  1. Fix SHA-256 snapshots of the two projects, excluding media, caches, and internal repository directories.

  2. Load the same task card in OpenCode 1.17.13, set the same permissions, 60-minute limit, and 65,536-token completion ceiling, and retain each model's native reasoning settings.

  3. Prohibit editing, network access, plugins, and subagents; allow only safe reading, retrieval, and directory checks.

  4. Use StackPerf to record every request, tool call, failure, token count, and cache event; anonymize the two final reports for blind evaluation, then perform factual verification of paths, symbols, and conclusions.

  5. Report scores, call efficiency, citation coverage, and unsupported claims separately; do not extrapolate a single 80/83 result into an overall model ranking.

Conclusions

This is a joint measurement of “model × harness × task,” not proof of Qwen3.8-Max's absolute capability. Qwen had advantages in architecture boundaries, replay, and fewer tool calls; Kimi had advantages in lifecycle and regeneration design. For high-value code review, having two models produce outputs in parallel and then cross-checking them is safer than sending one model's polished long report directly into implementation.

Limitations

  • The test used Qwen3.8-Max Preview; the article explicitly says to rerun it after changes to the GA version, official API, or open weights.

  • Each model had only one complete service path, one task, and one 60-minute session, so the effects of model capability cannot be separated from the endpoint, cache, rate limit, and harness.

  • The blind-evaluation score included human factual deductions, and the article does not disclose the complete task prompt, all tool schemas, or the raw turn-by-turn transcript.

  • The task concerned video-system architecture and codebase reading; it cannot replace SWE-bench, Terminal-Bench, Chinese-language, or multimodal regression tests.

Original evidence and data

The article makes public links to the two projects and the StackPerf repository, the 269-file scale, OpenCode version, permissions, time/token limits, model reasoning configuration, request/tool counts, citation counts, and factual-verification results; these fields are sufficient to reproduce the experiment's structure, but not the exact score of 80 without the same snapshots and service access.

Source excerpt or observation (short quote for compliance only)

The article describes Qwen's configuration as “enable_thinking: true and no fixed thinking budget”; this demonstrates the test configuration only and does not represent the default settings of every Qwen3.8-Max client.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Qwen3.8 Max

Use and compare models in Tabbit

Qwen3.8 Max

Related reviews

MediaOfficial Qwen Blog2026-08-03

Qwen3.8-Max: Official Release Notes and Complete Performance Results

MediaArtificial Analysis

Qwen3.8-Max: Artificial Analysis's Independent Index for Quality, Cost, Speed, and Verbosity

MediaNYU Shanghai RITS

Qwen3.8-Max: NYU Shanghai RITS Review of Agentic Index Evolution, Turns, and Hallucination Cost

MediaBenchLM2026-08-17

Qwen3.8 Max: BenchLM's Source-Verifiable Benchmark Ledger

Qwen3.8 Max

Related prompts

Mediaqwen.ai2026-08-03

Qwen3.8-Max: Reasoning Effort, Context Retention, and Agent Integration Prompting Guide

MediaEvoLink.AI Blog2026-08-03

Qwen3.8-Max Production Routing: EvoLink Prompting, Thinking Streams, and Tool Calling Guide

CommunityReddit, r/QwenAI2026-08-09

Qwen Studio + MCP: Prompting Qwen3.8-Max to Access Local Files and Permission Boundaries

CommunityReddit, r/QwenAI2026-08-06

Qwen3.8-Max Multi-turn Conversations: A Context-Recovery Prompt Derived from “Forgetting Follow-ups” Reports