Kimi K2.7 Code is a serious pilot candidate for repository work, coding agents and multimodal development. The practical catch is not its headline parameter count: Moonshot's official table charges $4 per million output tokens, while the high-speed variant charges $8. Forced thinking and preserved reasoning make output budget part of the architecture decision.
The decision anchor is the live Moonshot pricing table, checked September 20, 2026: K2.7 Code is $0.19/M cached input, $0.95/M cache-miss input and $4/M output; kimi-k2.7-code-highspeed doubles those lines. That is API billing, not a Kimi membership, GitHub Copilot usage price, provider rate or Tabbit subscription. (Kimi API guide)
Key takeaways
K2.7 Code is Moonshot's coding-focused route with 256K context, text/image/video input, thinking mode and OpenAI/Anthropic-compatible APIs.
The official model card reports a 1T-total/32B-active MoE, MoonViT vision encoder, native INT4 and approximately 30% lower thinking-token use than K2.6.
Moonshot's official benchmark table reports Kimi Code Bench v2 62.0, Program Bench 53.6, MLS Bench Lite 35.1, Kimi Claw 46.9 and MCP Mark Verified 81.1. These are vendor-reported evaluation rows.
Unsiloed's two-task comparison scored Kimi 53/60 on a FastAPI implementation versus GLM 48/60, but gave GLM deeper Saleor repository analysis. That is useful task evidence, not a general success rate.
GitHub Copilot hosts K2.7 Code on Azure and bills through its own usage-based model policy. Business and Enterprise administrators must enable it.
No Tabbit K2.7 Code task or screenshot was completed for this draft. Start with the Kimi K2.7 Code model resource.
What Kimi K2.7 Code actually is
Kimi K2.7 Code is the coding-specialized sibling in Moonshot's K2 family, not the same route as K2.6 general-purpose or K3 flagship. The prompt library and review collection preserve model-specific evidence.
| Question | Current snapshot | Boundary |
|---|---|---|
| Model ID | kimi-k2.7-code | Highspeed is a separate ID and price line. |
| Architecture | 1T total / 32B active MoE, 384 experts, 8 selected per token | Model-card specification, not a production cost guarantee. |
| Context | 262,144 tokens in official pricing table; model card says 256K | Treat these as the same approximate family boundary, not a larger 1M window. |
| Input | Text, image and video | Official API supports video; third-party deployments may not. |
| Reasoning | Thinking forced; preserve-thinking required | Clients must carry reasoning fields correctly. |
| API price | $0.19 cached input / $0.95 miss / $4 output per million | Highspeed: $0.38/$1.90/$8. Taxes and provider fees excluded. |
| Deployment | Moonshot API, HF weights, vLLM/SGLang/KTransformers, providers | Hardware, license and client support remain separate. |
K2.6, K2.7 Code and K3 are different choices
| Route | Best starting hypothesis | Boundary |
|---|---|---|
| K2.6 | General chat, multimodal agents, thinking or non-thinking | Same 256K family context, separate $0.16/$0.95/$4 API lines. |
| K2.7 Code | Repository implementation, debugging, coding agents and MCP | Thinking is forced; output can dominate cost. |
| K3 | New flagship for long-horizon coding and knowledge work | K3 has a 1M context and $3/$15 output economics; do not use K3 pricing for K2.7. |
The current quickstart recommends K3 as the default and K2.7 Code for coding-focused scenarios, with highspeed when output speed matters. That is a first-party routing recommendation, not a claim that K2.7 wins every repository task. The Kimi K3 pricing guide covers the separate K3 budget.
What the evidence actually supports
The model card's rows are internally coherent but vendor-reported: K2.7 Code rises from K2.6 on Kimi Code Bench v2 (50.9 to 62.0), Program Bench (48.3 to 53.6), MLS Bench Lite (26.7 to 35.1), Kimi Claw (42.9 to 46.9) and MCP Mark Verified (72.8 to 81.1). They describe the intended coding/agent surface, not an independent pass rate.
Unsiloed used the same FastAPI prompt and scored Kimi 53/60 versus GLM 48/60, then found GLM deeper on Saleor architecture and technical debt. The study's two tasks, undisclosed repeats and incomplete artifacts limit generalization. Agentic reasoning guidance is a better frame than “beats model X.”
How to get it without mixing products
Moonshot API: create an API key, choose
kimi-k2.7-code, record cache hit/miss, reasoning fields, tool calls and output tokens.Self-hosting: use the modified-MIT model card with vLLM, SGLang or KTransformers; hardware and video support must be tested locally.
Provider route: OpenRouter and inference providers can apply their own rate, fallback, data policy and limits.
Kimi Code or membership: consumer allowances and CLI quotas are not API billing.
GitHub Copilot: Azure-hosted, usage-based, gradual rollout; Business/Enterprise needs an administrator policy.
Tabbit Browser: a separate browser subscription and model-picker route. This article did not verify account-level availability.
Unknown risks and scenario self-check
Output economics: $4/M output is over four times cache-miss input; long reasoning, retries and preserved traces can dominate total cost.
Field compatibility: forced thinking and
reasoning_contentcan break clients or gateways that drop the field.Benchmark ownership: Kimi Code Bench, Kimi Claw and MCP rows are not independent production rates.
Repository fit: the two-task study found implementation and repository-analysis strengths split across models.
Policy/access: Copilot rollout, Kimi membership quotas, provider availability and Tabbit model selectors change independently.
| Your work | First move | Do not infer |
|---|---|---|
| Repository implementation | Fixed repo, tools, tests and a stop rule | Vendor coding score is not your pass rate. |
| Large architecture review | Compare K2.7 Code with a general reasoning model | Coding specialization guarantees no repository insight. |
| MCP agent | Disposable branch and least-privilege tools | High MCP score makes tool use safe. |
| Budget-sensitive coding | Compare cached context and output tokens | API price equals Kimi Code or Copilot subscription price. |
| Browser work | Read the agentic browser guide and check permissions | Model access grants no website authentication. |
What users report
The Kimi model-selection thread is unusually useful because the author explicitly says it is a docs breakdown, not a hands-on benchmark. Comments report K2.6 feeling better for coding, K2.7 consuming usage faster, Kimi Agent Swarm being useful for deep dives, and K2.7 being “plain stupid.” These observations have no shared fixture. Search also surfaced users describing K2.7 as near Claude/GPT quality but constrained by subscription usage; that is product feedback, not a model score. X exposed no stable post text and YouTube exposed test titles/chapters only.
A practical next step
Choose one reversible repository task: a small feature with tests, a bug fix with a known failing case, or an MCP workflow in a disposable branch. Record exact model ID, highspeed flag, cache state, reasoning/preserve-thinking fields, input/output tokens, tools, latency, retries, provider and human corrections. Compare it with K2.6 or your current coding model. Keep K2.7 Code only when accepted output offsets the $4/M output and review burden.
For browser workflows, use browser automation and Tabbit Browser as the product layer. Compare broader AI browser choices and Tabbit pricing; neither changes Moonshot API billing.
Verdict
Kimi K2.7 Code earns a controlled pilot for repository implementation, coding agents and MCP workflows. Its official API is inexpensive on input but not on output: $4/M, or $8/M for highspeed, with forced thinking and preserved reasoning. Vendor benchmarks show the intended strengths; a small independent comparison shows that implementation and repository analysis can diverge. Pin the route, budget output, isolate tools and let accepted work decide.
Sources
FAQ
What is Kimi K2.7 Code?
Kimi K2.7 Code is Moonshot's coding-focused agent model with a 256K context, text/image/video input, forced thinking and OpenAI/Anthropic-compatible API routes.
How much does Kimi K2.7 Code cost?
Moonshot's current API table lists $0.19 per million cached input tokens, $0.95 cache-miss input, and $4 output. The high-speed variant lists $0.38/$1.90/$8. Taxes and provider routes are separate.
What changed from Kimi K2.6?
The official model card describes K2.7 Code as coding-specialized, with 1T total/32B active parameters, MoonViT input, native INT4 and roughly 30% lower thinking-token use than K2.6. These are model-card claims, not an independent success rate.
Where can I access Kimi K2.7 Code?
Use the Moonshot Kimi API, self-host from the Hugging Face model card, supported inference providers, Kimi Code, or GitHub Copilot where rollout and policy allow. Each route has separate billing and limits.
Is Kimi K2.7 Code good for coding agents?
It merits a controlled repository and MCP pilot. The model card reports strong coding and tool scores, while a two-task independent comparison favored it for implementation but GLM 5.2 for repository analysis. Neither proves a universal win.
Can I use Kimi K2.7 Code in Tabbit?
This overview did not run an account-level Tabbit test. Check the live picker and run a reversible coding or research task before planning around it.