Hugging Face / Moonshot AI Official Model CardVendor report
Kimi K2.7 Code: Official Hugging Face Model Specifications and Full Benchmark Data
Kimi K2.7 Code is a long-horizon coding and agent-specialized model built on an MoE architecture (1T total parameters / 32B active) , natively integrating the MoonViT multimodal vision encoder and out-of-the-box INT4 quantization, achieving a massive leap in coding performance while cutting thinking token consumption by roughly 30% compared to K2.6.
- Evidence
- Vendor report
- Boundary
- Several benchmarks in the official comparison table are internally developed evaluation suites by Moonshot (such as Kimi Code Bench v2 and Kimi Claw 24/7) , which require cross-validation against open-source third-party benchmarks.
Unsiloed AI Engineering Blog / Reddit r/LangChainIndependent measurement
Unsiloed Benchmark: Kimi K2.7 Code vs GLM 5.2 Controlled Benchmark on Real-World Code Generation and Large Repository Analysis
In strictly controlled tests using identical prompts, Kimi K2.7 beat GLM 5.2 (48/60) with a score of 53/60 in scaffolding a runnable greenfield project (FastAPI) thanks to complete components and zero missing dependencies; meanwhile, in deconstructing a massive repository (Saleor) end-to-end, GLM 5.2 came out on top by leveraging its 1M context window to unearth deeper implementation details.
- Evidence
- Independent measurement
- Boundary
- This benchmark relied on single-turn zero-shot / few-shot prompt comparisons and did not evaluate final convergence performance in multi-turn Agent self-correction loops (e.g., self-running pytest to resolve missing components) .
Reddit r/windsurf / Devin.ai (Cognition)Independent measurement
Devin Team: FrontierCode Extended Benchmark and Long-Horizon Engineering Performance
On the independent FrontierCode Extended benchmark built by the Devin team for real-world software engineering tasks, Kimi K2.7 Code achieved a 39.5% pass rate, placing it firmly in the competitive tier alongside top-tier proprietary models. It excels at generating standalone UI components and self-contained features, but remains constrained by its context window and memory span during long-sequence multi-file refactoring.
- Evidence
- Independent measurement
- Boundary
- FrontierCode Extended incorporates Devin platform-specific agent toolsets and execution feedback mechanisms; switching to alternative agent frameworks (such as SWE-agent or Aider) may produce different pass rates.