Kimi K2.7 Code

Kimi K2.7 Code · Reviews and evidence

Which Kimi K2.7 Code conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

Kimi K2.7 Code is a long-horizon coding and agent-specialized model built on an MoE architecture (1T total parameters / 32B active) , natively integrating the MoonViT multimodal vision encoder and out-of-the-box INT4 quantization, achieving a massive leap in coding performance while cutting thinking token consumption by roughly 30% compared to K2.6.

Hugging Face / Moonshot AI Official Model Card · Read evidence

In strictly controlled tests using identical prompts, Kimi K2.7 beat GLM 5.2 (48/60) with a score of 53/60 in scaffolding a runnable greenfield project (FastAPI) thanks to complete components and zero missing dependencies; meanwhile, in deconstructing a massive repository (Saleor) end-to-end, GLM 5.2 came out on top by leveraging its 1M context window to unearth deeper implementation details.

Unsiloed AI Engineering Blog / Reddit r/LangChain · Read evidence

On the independent FrontierCode Extended benchmark built by the Devin team for real-world software engineering tasks, Kimi K2.7 Code achieved a 39.5% pass rate, placing it firmly in the competitive tier alongside top-tier proprietary models. It excels at generating standalone UI components and self-contained features, but remains constrained by its context window and memory span during long-sequence multi-file refactoring.

Reddit r/windsurf / Devin.ai (Cognition) · Read evidence

Full reviews and related reading

Read the full analysis

Overview · English

Kimi K2.7 Code: What It Is, What It Costs, and Who It Fits

A sourced Kimi K2.7 Code overview: the $4/M output anchor, 256K multimodal coding model, K2.6/K3 boundary, access routes and pilot risks.

Selected evidence

Media / benchmarkVendor report

Kimi K2.7 Code: Official Hugging Face Model Specifications and Full Benchmark Data

Kimi K2.7 Code is a long-horizon coding and agent-specialized model built on an MoE architecture (1T total parameters / 32B active) , natively integrating the MoonViT multimodal vision encoder and out-of-the-box INT4 quantization, achieving a massive leap in coding performance while cutting thinking token consumption by roughly 30% compared to K2.6.

SourceHugging Face / Moonshot AI Official Model Card
Published2026-06-12
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-06-12.
Harness/task
Base architecture: MoE architecture with 61 layers in total (including 1 dense layer) , 384 experts, activating 8 routed experts + 1 shared expert per token.; Attention and activation: Multi-Head Latent Attention (MLA) mechanism, SwiGLU activation function, 160K vocabulary size, and 256K context window.
Sample/gaps
Limitations noted: Several benchmarks in the official comparison table are internally developed evaluation suites by Moonshot (such as Kimi Code Bench v2 and Kimi Claw 24/7) , which require cross-validation against open-source third-party benchmarks.; The mandatory requirement to preserve `reasoningcontent` introduces compatibility barriers for third-party clients and API gateways that do not support reasoning field round-tripping.
CodingAgent
CommunityIndependent measurement

Unsiloed Benchmark: Kimi K2.7 Code vs GLM 5.2 Controlled Benchmark on Real-World Code Generation and Large Repository Analysis

In strictly controlled tests using identical prompts, Kimi K2.7 beat GLM 5.2 (48/60) with a score of 53/60 in scaffolding a runnable greenfield project (FastAPI) thanks to complete components and zero missing dependencies; meanwhile, in deconstructing a massive repository (Saleor) end-to-end, GLM 5.2 came out on top by leveraging its 1M context window to unearth deeper implementation details.

SourceUnsiloed AI Engineering Blog / Reddit r/LangChain
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-07-20.
Harness/task
Kimi K2.7 Code: MoE 1T total parameters / 32B active, 256K context window, official API pricing at $0.95 input ($0.19 cached) / $4.00 output per 1M tokens.; GLM 5.2: MoE 744B–753B total parameters / 40B active, 1M context window, official API pricing at $1.40 input ($0.26 cached) / $4.40 output per 1M tokens.
Sample/gaps
Limitations noted: This benchmark relied on single-turn zero-shot / few-shot prompt comparisons and did not evaluate final convergence performance in multi-turn Agent self-correction loops (e.g., self-running pytest to resolve missing components) .; Pricing comparisons are based solely on official standard API rates and do not account for third-party aggregators or specific subscription plan discounts.
CodingAgent
CommunityIndependent measurement

Devin Team: FrontierCode Extended Benchmark and Long-Horizon Engineering Performance

On the independent FrontierCode Extended benchmark built by the Devin team for real-world software engineering tasks, Kimi K2.7 Code achieved a 39.5% pass rate, placing it firmly in the competitive tier alongside top-tier proprietary models. It excels at generating standalone UI components and self-contained features, but remains constrained by its context window and memory span during long-sequence multi-file refactoring.

SourceReddit r/windsurf / Devin.ai (Cognition)
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-06-24.
Harness/task
Evaluation Benchmark: FrontierCode Extended (a comprehensive benchmark suite by the Devin / Cognition team designed to evaluate real-world end-to-end software engineering tasks) .; Execution Environment: Real agent execution environments across Devin Desktop and Devin CLI.
Sample/gaps
Limitations noted: FrontierCode Extended incorporates Devin platform-specific agent toolsets and execution feedback mechanisms; switching to alternative agent frameworks (such as SWE-agent or Aider) may produce different pass rates.; The client-side limit of a 200K context during testing prevented the model from fully leveraging its native 256K token potential.
CodingAgent
CommunityPersonal experience

OpenCode Community: Real-World Agentic Coding Cost & Tool Loop Efficiency Comparison

Nominal Unit Price $\neq$ Real-World Agent Cost: In autonomous agent environments, if a model lacks precise tool-calling decision capabilities, it easily falls into a death loop of "repeated file reads $\to$ repeated failed command executions $\to$ lengthy retries," causing context to explode and token consumption to spike geometrically.

SourceReddit r/opencode
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-08-04.
Harness/task
Nominal Unit Price $\neq$ Real-World Agent Cost: In autonomous agent environments, if a model lacks precise tool-calling decision capabilities, it easily falls into a death loop of "repeated file reads $\to$ repeated failed command executions $\to$ lengthy retries," causing context to explode and token consumption to spike geometrically.
Sample/gaps
Limitations noted: Cost data is heavily influenced by the agent framework's prompt design, system guardrails (Loop Detection) , and context truncation strategies; in advanced harnesses equipped with strict deduplication and loop interception, the cost gap between the two may narrow.; Data originates from community real-world development usage statistics and LiveBench aggregate benchmarks, which carry sample distribution variations.
CodingAgent

All sources

All sources

6 / 6
Media / benchmarkVendor report

Kimi K2.7 Code: Official Hugging Face Model Specifications and Full Benchmark Data

Kimi K2.7 Code is a long-horizon coding and agent-specialized model built on an MoE architecture (1T total parameters / 32B active) , natively integrating the MoonViT multimodal vision encoder and out-of-the-box INT4 quantization, achieving a massive leap in coding performance while cutting thinking token consumption by roughly 30% compared to K2.6.

SourceHugging Face / Moonshot AI Official Model Card
Published2026-06-12
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-06-12.
Harness/task
Base architecture: MoE architecture with 61 layers in total (including 1 dense layer) , 384 experts, activating 8 routed experts + 1 shared expert per token.; Attention and activation: Multi-Head Latent Attention (MLA) mechanism, SwiGLU activation function, 160K vocabulary size, and 256K context window.
Sample/gaps
Limitations noted: Several benchmarks in the official comparison table are internally developed evaluation suites by Moonshot (such as Kimi Code Bench v2 and Kimi Claw 24/7) , which require cross-validation against open-source third-party benchmarks.; The mandatory requirement to preserve `reasoningcontent` introduces compatibility barriers for third-party clients and API gateways that do not support reasoning field round-tripping.
CodingAgent
CommunityIndependent measurement

Unsiloed Benchmark: Kimi K2.7 Code vs GLM 5.2 Controlled Benchmark on Real-World Code Generation and Large Repository Analysis

In strictly controlled tests using identical prompts, Kimi K2.7 beat GLM 5.2 (48/60) with a score of 53/60 in scaffolding a runnable greenfield project (FastAPI) thanks to complete components and zero missing dependencies; meanwhile, in deconstructing a massive repository (Saleor) end-to-end, GLM 5.2 came out on top by leveraging its 1M context window to unearth deeper implementation details.

SourceUnsiloed AI Engineering Blog / Reddit r/LangChain
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-07-20.
Harness/task
Kimi K2.7 Code: MoE 1T total parameters / 32B active, 256K context window, official API pricing at $0.95 input ($0.19 cached) / $4.00 output per 1M tokens.; GLM 5.2: MoE 744B–753B total parameters / 40B active, 1M context window, official API pricing at $1.40 input ($0.26 cached) / $4.40 output per 1M tokens.
Sample/gaps
Limitations noted: This benchmark relied on single-turn zero-shot / few-shot prompt comparisons and did not evaluate final convergence performance in multi-turn Agent self-correction loops (e.g., self-running pytest to resolve missing components) .; Pricing comparisons are based solely on official standard API rates and do not account for third-party aggregators or specific subscription plan discounts.
CodingAgent
CommunityIndependent measurement

Devin Team: FrontierCode Extended Benchmark and Long-Horizon Engineering Performance

On the independent FrontierCode Extended benchmark built by the Devin team for real-world software engineering tasks, Kimi K2.7 Code achieved a 39.5% pass rate, placing it firmly in the competitive tier alongside top-tier proprietary models. It excels at generating standalone UI components and self-contained features, but remains constrained by its context window and memory span during long-sequence multi-file refactoring.

SourceReddit r/windsurf / Devin.ai (Cognition)
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-06-24.
Harness/task
Evaluation Benchmark: FrontierCode Extended (a comprehensive benchmark suite by the Devin / Cognition team designed to evaluate real-world end-to-end software engineering tasks) .; Execution Environment: Real agent execution environments across Devin Desktop and Devin CLI.
Sample/gaps
Limitations noted: FrontierCode Extended incorporates Devin platform-specific agent toolsets and execution feedback mechanisms; switching to alternative agent frameworks (such as SWE-agent or Aider) may produce different pass rates.; The client-side limit of a 200K context during testing prevented the model from fully leveraging its native 256K token potential.
CodingAgent
CommunityPersonal experience

OpenCode Community: Real-World Agentic Coding Cost & Tool Loop Efficiency Comparison

Nominal Unit Price $\neq$ Real-World Agent Cost: In autonomous agent environments, if a model lacks precise tool-calling decision capabilities, it easily falls into a death loop of "repeated file reads $\to$ repeated failed command executions $\to$ lengthy retries," causing context to explode and token consumption to spike geometrically.

SourceReddit r/opencode
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-08-04.
Harness/task
Nominal Unit Price $\neq$ Real-World Agent Cost: In autonomous agent environments, if a model lacks precise tool-calling decision capabilities, it easily falls into a death loop of "repeated file reads $\to$ repeated failed command executions $\to$ lengthy retries," causing context to explode and token consumption to spike geometrically.
Sample/gaps
Limitations noted: Cost data is heavily influenced by the agent framework's prompt design, system guardrails (Loop Detection) , and context truncation strategies; in advanced harnesses equipped with strict deduplication and loop interception, the cost gap between the two may narrow.; Data originates from community real-world development usage statistics and LiveBench aggregate benchmarks, which carry sample distribution variations.
CodingAgent
CommunityPersonal experience

Reddit Community: Where to Draw the Line Between Kimi K2.7 Code, K2.6, and K2.5

The reusable value of this post is that it establishes a model-division hypothesis, rather than proving that K2.7 Code wins every real-world task: let the coding-specialized model handle repository tasks, K2.6 handle general-purpose multimodal agents, and K2.5 handle low-cost ordinary work, while using caching to control the cost of repeated context.

SourceReddit, r/kimi
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-08-18.
Harness/task
The official API prices recorded in the article are: K2.7 Code cache-hit input/miss input/output at $0.19/$0.95/$4.00 per million tokens; K2.6 at $0.16/$0.95/$4.00; and K2.5 at $0.10/$0.60/$3.00.; The author provides no model run scores and explicitly says they are looking to collect real-world day-to-day usage feedback.
Sample/gaps
Limitations noted: “K2.7 Code is better suited to complete repositories” is a routing recommendation, not a measured pass rate, latency result, or error sample.; “Output tokens cost more” must not be interpreted as a uniform billing rule across all providers.
CodingAgent
CommunityPersonal experience

Reddit Community: Harness Integration Pitfalls and Reasoning Token Mishandling Hands-on Analysis

Strict Protocol Constraints of K2.7: Kimi K2.7 strictly enforces `thinking=enabled` and requires that `reasoningcontent` be completely preserved across multi-turn tool interactions. If a third-party harness drops the assistant's thinking content or converts it to plain text, the model loses its prior reasoning context, directly causing logical disconnects and repetitive tool invocations.

SourceReddit r/kimi
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Model/version
Kimi-K2.7-Code; source date: 2026-06-23.
Harness/task
Problematic Client Environments: Allegreto, custom simple Agent loops, and generic proxy gateways that have not adapted to Kimi's chain-of-thought pass-back protocol.; Tech Stacks Involved: React frontend, C backend projects.
Sample/gaps
Limitations noted: This post reflects real-world troubleshooting logs from the early post-launch period when the third-party ecosystem had not fully adapted to Kimi's new protocol. While highly valuable as a guide for avoiding pitfalls, it does not represent the model's true upper-bound performance in standard environments.
CodingAgent

Kimi K2.7 Code

Compare Kimi K2.7 Code in Tabbit

Model access, features, and permissions depend on your current client account.