TabbitBlog

Kimi K2.7 Code: What It Is, What It Costs, and Who It Fits

A sourced Kimi K2.7 Code overview: the $4/M output anchor, 256K multimodal coding model, K2.6/K3 boundary, access routes and pilot risks.

In this article
  1. Key takeaways
  2. What Kimi K2.7 Code actually is
  3. K2.6, K2.7 Code and K3 are different choices
  4. What the evidence actually supports
  5. How to get it without mixing products
  6. Unknown risks and scenario self-check
  7. What users report
  8. A practical next step
  9. Verdict
  10. Sources

Kimi K2.7 Code is a serious pilot candidate for repository work, coding agents and multimodal development. The practical catch is not its headline parameter count: Moonshot's official table charges $4 per million output tokens, while the high-speed variant charges $8. Forced thinking and preserved reasoning make output budget part of the architecture decision.

The decision anchor is the live Moonshot pricing table, checked September 20, 2026: K2.7 Code is $0.19/M cached input, $0.95/M cache-miss input and $4/M output; kimi-k2.7-code-highspeed doubles those lines. That is API billing, not a Kimi membership, GitHub Copilot usage price, provider rate or Tabbit subscription. (Kimi API guide)

Key takeaways

  • K2.7 Code is Moonshot's coding-focused route with 256K context, text/image/video input, thinking mode and OpenAI/Anthropic-compatible APIs.

  • The official model card reports a 1T-total/32B-active MoE, MoonViT vision encoder, native INT4 and approximately 30% lower thinking-token use than K2.6.

  • Moonshot's official benchmark table reports Kimi Code Bench v2 62.0, Program Bench 53.6, MLS Bench Lite 35.1, Kimi Claw 46.9 and MCP Mark Verified 81.1. These are vendor-reported evaluation rows.

  • Unsiloed's two-task comparison scored Kimi 53/60 on a FastAPI implementation versus GLM 48/60, but gave GLM deeper Saleor repository analysis. That is useful task evidence, not a general success rate.

  • GitHub Copilot hosts K2.7 Code on Azure and bills through its own usage-based model policy. Business and Enterprise administrators must enable it.

  • No Tabbit K2.7 Code task or screenshot was completed for this draft. Start with the Kimi K2.7 Code model resource.

What Kimi K2.7 Code actually is

Kimi K2.7 Code is the coding-specialized sibling in Moonshot's K2 family, not the same route as K2.6 general-purpose or K3 flagship. The prompt library and review collection preserve model-specific evidence.

QuestionCurrent snapshotBoundary
Model IDkimi-k2.7-codeHighspeed is a separate ID and price line.
Architecture1T total / 32B active MoE, 384 experts, 8 selected per tokenModel-card specification, not a production cost guarantee.
Context262,144 tokens in official pricing table; model card says 256KTreat these as the same approximate family boundary, not a larger 1M window.
InputText, image and videoOfficial API supports video; third-party deployments may not.
ReasoningThinking forced; preserve-thinking requiredClients must carry reasoning fields correctly.
API price$0.19 cached input / $0.95 miss / $4 output per millionHighspeed: $0.38/$1.90/$8. Taxes and provider fees excluded.
DeploymentMoonshot API, HF weights, vLLM/SGLang/KTransformers, providersHardware, license and client support remain separate.

K2.6, K2.7 Code and K3 are different choices

RouteBest starting hypothesisBoundary
K2.6General chat, multimodal agents, thinking or non-thinkingSame 256K family context, separate $0.16/$0.95/$4 API lines.
K2.7 CodeRepository implementation, debugging, coding agents and MCPThinking is forced; output can dominate cost.
K3New flagship for long-horizon coding and knowledge workK3 has a 1M context and $3/$15 output economics; do not use K3 pricing for K2.7.

The current quickstart recommends K3 as the default and K2.7 Code for coding-focused scenarios, with highspeed when output speed matters. That is a first-party routing recommendation, not a claim that K2.7 wins every repository task. The Kimi K3 pricing guide covers the separate K3 budget.

What the evidence actually supports

The model card's rows are internally coherent but vendor-reported: K2.7 Code rises from K2.6 on Kimi Code Bench v2 (50.9 to 62.0), Program Bench (48.3 to 53.6), MLS Bench Lite (26.7 to 35.1), Kimi Claw (42.9 to 46.9) and MCP Mark Verified (72.8 to 81.1). They describe the intended coding/agent surface, not an independent pass rate.

Unsiloed used the same FastAPI prompt and scored Kimi 53/60 versus GLM 48/60, then found GLM deeper on Saleor architecture and technical debt. The study's two tasks, undisclosed repeats and incomplete artifacts limit generalization. Agentic reasoning guidance is a better frame than “beats model X.”

How to get it without mixing products

  1. Moonshot API: create an API key, choose kimi-k2.7-code, record cache hit/miss, reasoning fields, tool calls and output tokens.

  2. Self-hosting: use the modified-MIT model card with vLLM, SGLang or KTransformers; hardware and video support must be tested locally.

  3. Provider route: OpenRouter and inference providers can apply their own rate, fallback, data policy and limits.

  4. Kimi Code or membership: consumer allowances and CLI quotas are not API billing.

  5. GitHub Copilot: Azure-hosted, usage-based, gradual rollout; Business/Enterprise needs an administrator policy.

  6. Tabbit Browser: a separate browser subscription and model-picker route. This article did not verify account-level availability.

Unknown risks and scenario self-check

  • Output economics: $4/M output is over four times cache-miss input; long reasoning, retries and preserved traces can dominate total cost.

  • Field compatibility: forced thinking and reasoning_content can break clients or gateways that drop the field.

  • Benchmark ownership: Kimi Code Bench, Kimi Claw and MCP rows are not independent production rates.

  • Repository fit: the two-task study found implementation and repository-analysis strengths split across models.

  • Policy/access: Copilot rollout, Kimi membership quotas, provider availability and Tabbit model selectors change independently.

Your workFirst moveDo not infer
Repository implementationFixed repo, tools, tests and a stop ruleVendor coding score is not your pass rate.
Large architecture reviewCompare K2.7 Code with a general reasoning modelCoding specialization guarantees no repository insight.
MCP agentDisposable branch and least-privilege toolsHigh MCP score makes tool use safe.
Budget-sensitive codingCompare cached context and output tokensAPI price equals Kimi Code or Copilot subscription price.
Browser workRead the agentic browser guide and check permissionsModel access grants no website authentication.

What users report

The Kimi model-selection thread is unusually useful because the author explicitly says it is a docs breakdown, not a hands-on benchmark. Comments report K2.6 feeling better for coding, K2.7 consuming usage faster, Kimi Agent Swarm being useful for deep dives, and K2.7 being “plain stupid.” These observations have no shared fixture. Search also surfaced users describing K2.7 as near Claude/GPT quality but constrained by subscription usage; that is product feedback, not a model score. X exposed no stable post text and YouTube exposed test titles/chapters only.

A practical next step

Choose one reversible repository task: a small feature with tests, a bug fix with a known failing case, or an MCP workflow in a disposable branch. Record exact model ID, highspeed flag, cache state, reasoning/preserve-thinking fields, input/output tokens, tools, latency, retries, provider and human corrections. Compare it with K2.6 or your current coding model. Keep K2.7 Code only when accepted output offsets the $4/M output and review burden.

For browser workflows, use browser automation and Tabbit Browser as the product layer. Compare broader AI browser choices and Tabbit pricing; neither changes Moonshot API billing.

Tabbit Browser

Verdict

Kimi K2.7 Code earns a controlled pilot for repository implementation, coding agents and MCP workflows. Its official API is inexpensive on input but not on output: $4/M, or $8/M for highspeed, with forced thinking and preserved reasoning. Vendor benchmarks show the intended strengths; a small independent comparison shows that implementation and repository analysis can diverge. Pin the route, budget output, isolate tools and let accepted work decide.

Sources

FAQ

What is Kimi K2.7 Code?

Kimi K2.7 Code is Moonshot's coding-focused agent model with a 256K context, text/image/video input, forced thinking and OpenAI/Anthropic-compatible API routes.

How much does Kimi K2.7 Code cost?

Moonshot's current API table lists $0.19 per million cached input tokens, $0.95 cache-miss input, and $4 output. The high-speed variant lists $0.38/$1.90/$8. Taxes and provider routes are separate.

What changed from Kimi K2.6?

The official model card describes K2.7 Code as coding-specialized, with 1T total/32B active parameters, MoonViT input, native INT4 and roughly 30% lower thinking-token use than K2.6. These are model-card claims, not an independent success rate.

Where can I access Kimi K2.7 Code?

Use the Moonshot Kimi API, self-host from the Hugging Face model card, supported inference providers, Kimi Code, or GitHub Copilot where rollout and policy allow. Each route has separate billing and limits.

Is Kimi K2.7 Code good for coding agents?

It merits a controlled repository and MCP pilot. The model card reports strong coding and tool scores, while a two-task independent comparison favored it for implementation but GLM 5.2 for repository analysis. Neither proves a universal win.

Can I use Kimi K2.7 Code in Tabbit?

This overview did not run an account-level Tabbit test. Check the live picker and run a reversible coding or research task before planning around it.

Take the next step

Let Tabbit work alongside you.

Research across tabs, automate repetitive browser work, and keep every piece of context within reach.