On the same 269-file, 60-minute StackPerf codebase architecture task using OpenCode 1.17.13, Qwen3.8-Max Preview scored 80 and Kimi K3 scored 83; Qwen was stronger on system boundaries, evidence citations, and replay metadata, while Kimi was stronger on revision, regeneration, and scenario lifecycle, and both still required independent factual verification.
Suitable tasks: Long-context architecture reviews that require reading an unfamiliar codebase, comparing system responsibilities, designing data contracts, writing migration plans, and maintaining an evidence ledger.
Unsuitable tasks: Treating one codebase architecture blind test as a general coding benchmark, single-turn Q&A, or the current score for Qwen3.8-Max GA.
Applicable model versions: Qwen3.8-Max-Preview (the 2026-07-19 service snapshot), not a direct substitute for later GA versions.
Applicable clients, Agents, or APIs: OpenCode 1.17.13; Qwen used Alibaba's international Token Plan endpoint, and Kimi used the Kimi Code subscription endpoint.
Recommended reasoning tier and parameters: Qwen enable_thinking: true, with no fixed thinking budget; the article does not disclose temperature or the complete system prompt, which should not be filled in speculatively.
The input consisted of read-only snapshots of two unfamiliar projects, trilogy-group/ttv-pipeline and kumanday/media-tooling, totaling 269 files; generated media, caches, virtual environments, and internal repository directories were excluded.
The task required file-by-file/line-by-line evidence analysis, distinguishing global planning from provider-specific generation, comparing 15/30-second video generation, designing typed/JSON contracts, and proposing migration phases, tests, risks, and an evidence ledger.
Both models used the same task card, permissions, and 60-minute wall-clock limit, with a completion ceiling of 65,536 tokens.
Read, search, listing, and safety checks were allowed; editing, network access, writing source files, plugins, and subagents were prohibited.
StackPerf registered the sessions and collected request, tool, and output metrics; before blind evaluation, the reports were renamed Report A/B, followed by a factual verification stage.
Qwen3.8-Max Preview: 80/100 after factual deductions; Kimi K3: 83/100.
Qwen: 22 gateway requests and 44 tool calls, all tool calls successful; the report was longer, with 354 citations, and verification found no fabricated paths or symbols, but 2 cited line numbers were inaccurate and 7 groups of conclusions exceeded the code evidence.
Kimi: 53 tool calls, with 2 compound shell commands rejected by policy before recovery; it had 274 citations, no fabricated paths or symbols were found, and it likewise had 7 groups of conclusions that exceeded the evidence.
Both models reached the same system boundary: one side owned global planning and final assembly, while the other handled provider-specific generation, retries, and provenance, connected through a versioned contract.
Qwen's GenerationManifest / GenerationResult and replay records were more complete, recording model, seed, prompt, references, provider request ID, media URI, duration, and cost; Kimi modeled revision, supersession, retry history, and take invalidation more completely.
The cache hit rate for repeated prompts on both provider paths exceeded 90%; this figure is affected by the provider, routing, and session cache.
Fix SHA-256 snapshots of the two projects, excluding media, caches, and internal repository directories.
Load the same task card in OpenCode 1.17.13, set the same permissions, 60-minute limit, and 65,536-token completion ceiling, and retain each model's native reasoning settings.
Prohibit editing, network access, plugins, and subagents; allow only safe reading, retrieval, and directory checks.
Use StackPerf to record every request, tool call, failure, token count, and cache event; anonymize the two final reports for blind evaluation, then perform factual verification of paths, symbols, and conclusions.
Report scores, call efficiency, citation coverage, and unsupported claims separately; do not extrapolate a single 80/83 result into an overall model ranking.
This is a joint measurement of “model × harness × task,” not proof of Qwen3.8-Max's absolute capability. Qwen had advantages in architecture boundaries, replay, and fewer tool calls; Kimi had advantages in lifecycle and regeneration design. For high-value code review, having two models produce outputs in parallel and then cross-checking them is safer than sending one model's polished long report directly into implementation.
The test used Qwen3.8-Max Preview; the article explicitly says to rerun it after changes to the GA version, official API, or open weights.
Each model had only one complete service path, one task, and one 60-minute session, so the effects of model capability cannot be separated from the endpoint, cache, rate limit, and harness.
The blind-evaluation score included human factual deductions, and the article does not disclose the complete task prompt, all tool schemas, or the raw turn-by-turn transcript.
The task concerned video-system architecture and codebase reading; it cannot replace SWE-bench, Terminal-Bench, Chinese-language, or multimodal regression tests.
The article makes public links to the two projects and the StackPerf repository, the 269-file scale, OpenCode version, permissions, time/token limits, model reasoning configuration, request/tool counts, citation counts, and factual-verification results; these fields are sufficient to reproduce the experiment's structure, but not the exact score of 80 without the same snapshots and service access.
The article describes Qwen's configuration as “enable_thinking: true and no fixed thinking budget”; this demonstrates the test configuration only and does not represent the default settings of every Qwen3.8-Max client.
Qwen3.8 Max