Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Reviews and evidence

Kimi K2.7 Code · Community source · Independent measurement

Devin Team: FrontierCode Extended Benchmark and Long-Horizon Engineering Performance

On the independent FrontierCode Extended benchmark built by the Devin team for real-world software engineering tasks, Kimi K2.7 Code achieved a 39.5% pass rate, placing it firmly in the competitive tier alongside top-tier proprietary models. It excels at generating standalone UI components and self-contained features, but remains constrained by its context window and memory span during long-sequence multi-file refactoring.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceIndependent measurementEdited 2026-09-20

Test conditions

Model/version
Kimi-K2.7-Code; source date: 2026-06-24.
Harness/task
Evaluation Benchmark: FrontierCode Extended (a comprehensive benchmark suite by the Devin / Cognition team designed to evaluate real-world end-to-end software engineering tasks) .; Execution Environment: Real agent execution environments across Devin Desktop and Devin CLI.
Sample/gaps
Limitations noted: FrontierCode Extended incorporates Devin platform-specific agent toolsets and execution feedback mechanisms; switching to alternative agent frameworks (such as SWE-agent or Aider) may produce different pass rates.; The client-side limit of a 200K context during testing prevented the model from fully leveraging its native 256K token potential.

Key data and applicable tasks

One-Sentence Takeaway

On the independent FrontierCode Extended benchmark built by the Devin team for real-world software engineering tasks, Kimi K2.7 Code achieved a 39.5% pass rate, placing it firmly in the competitive tier alongside top-tier proprietary models. It excels at generating standalone UI components and self-contained features, but remains constrained by its context window and memory span during long-sequence multi-file refactoring.

Test Environment, Input / Configuration

  • Evaluation Benchmark: FrontierCode Extended (a comprehensive benchmark suite by the Devin / Cognition team designed to evaluate real-world end-to-end software engineering tasks) .

  • Execution Environment: Real agent execution environments across Devin Desktop and Devin CLI.

  • Model Comparison Group: Claude Opus 4.8, GPT-5.5, GLM 5.2, Kimi K2.7 Code.

  • Client Constraints: During testing, the Devin client capped the allocated context window for K2.7 / GLM 5.2 at approximately 200K tokens.

Results Data

FrontierCode Extended official published results (Pass Rate) :

ModelModel TypeFrontierCode Extended Score
Claude Opus 4.8Frontier Proprietary Flagship51.8%
GPT-5.5Frontier Proprietary Flagship44.8%
GLM 5.2Open-Weight MoE43.0%
Kimi K2.7 CodeOpen-Weight MoE39.5%

Community developer feedback distribution in real-world engineering:

  • Self-Contained Features / UI Tasks: Kimi K2.7 generates code exceptionally fast with snappy responsiveness and high out-of-the-box usability.

  • **Large-Scale Long-Horizon Refactoring (>15 consecutive files modified) **: Performance degrades when the agent must maintain precise memory of changes in files 1–8 while working through files 9–15.

Conclusion

  1. Open-Source Models Enter the Top Tier: Kimi K2.7 Code (39.5%) and GLM 5.2 (43.0%) demonstrated problem-solving capabilities on real-world complex engineering tasks that approach top-tier proprietary models (GPT-5.5 at 44.8%) .

  2. Task Suitability Divide:

    • High-Win Scenarios: End-to-end single-feature development, standalone module refactoring, UI/frontend page implementation, and unit test generation.

    • Challenging Scenarios: Long-horizon cascading refactoring involving more than 15 interrelated files and synchronized dependency updates across multiple repositories.

Limitations

  • FrontierCode Extended incorporates Devin platform-specific agent toolsets and execution feedback mechanisms; switching to alternative agent frameworks (such as SWE-agent or Aider) may produce different pass rates.

  • The client-side limit of a 200K context during testing prevented the model from fully leveraging its native 256K token potential.

Reproduction Steps

  1. Select Kimi K2.7 Code as the underlying execution model in Devin Desktop / CLI.

  2. Configure two sets of engineering tasks: standalone feature creation and multi-file refactoring.

  3. Record the number of execution turns, tool invocation success rate, final build/test pass status, and token consumption.

Raw Evidence and Data

Devin team engineer theodormarcu published the specific percentage scores above in an official announcement, confirming FrontierCode Extended as the comparative benchmark.

Source Excerpts or Observations (Short Compliance Excerpts Only)

  • Official Announcement: "GLM 5.2 scores 43.0% and Kimi K2.7 scores 39.5% on FrontierCode Extended — placing them in a competitive tier alongside GPT-5.5 (44.8%) and Claude Opus 4.8 (51.8%) ."

  • Senior Community Developer Observation: "200k covers most single-feature work but shows up on large refactors — when you're touching 15 files in sequence and need the model to hold what changed in files 1-8 while working on 9-15, the ceiling matters."

What this supports

  • Open-Source Models Enter the Top Tier: Kimi K2.7 Code (39.5%) and GLM 5.2 (43.0%) demonstrated problem-solving capabilities on real-world complex engineering tasks that approach top-tier proprietary models (GPT-5.5 at 44.8%) .
  • High-Win Scenarios: End-to-end single-feature development, standalone module refactoring, UI/frontend page implementation, and unit test generation.

What this does not support

  • FrontierCode Extended incorporates Devin platform-specific agent toolsets and execution feedback mechanisms; switching to alternative agent frameworks (such as SWE-agent or Aider) may produce different pass rates.
  • The client-side limit of a 200K context during testing prevented the model from fully leveraging its native 256K token potential.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit r/windsurf / Devin.ai (Cognition) · theodormarcu (Devin / Windsurf Team Engineer) · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Kimi K2.7 Code

Compare Kimi K2.7 Code in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Kimi K2.7 Code: What It Is, What It Costs, and Who It Fits

A sourced Kimi K2.7 Code overview: the $4/M output anchor, 256K multimodal coding model, K2.6/K3 boundary, access routes and pilot risks.

Related reviews

Unsiloed Benchmark: Kimi K2.7 Code vs GLM 5.2 Controlled Benchmark on Real-World Code Generation and Large Repository AnalysisIn strictly controlled tests using identical prompts, Kimi K2.7 beat GLM 5.2 (48/60) with a score of 53/60 in scaffolding a runnable greenfield project (FastAPI) thanks to complete components and zero missing dependencies; meanwhile, in deconstructing a massive repository (Saleor) end-to-end, GLM 5.2 came out on top by leveraging its 1M context window to unearth deeper implementation details.OpenCode Community: Real-World Agentic Coding Cost & Tool Loop Efficiency ComparisonNominal Unit Price $\neq$ Real-World Agent Cost: In autonomous agent environments, if a model lacks precise tool-calling decision capabilities, it easily falls into a death loop of "repeated file reads $\to$ repeated failed command executions $\to$ lengthy retries," causing context to explode and token consumption to spike geometrically.Kimi K2.7 Code: Official Hugging Face Model Specifications and Full Benchmark DataKimi K2.7 Code is a long-horizon coding and agent-specialized model built on an MoE architecture (1T total parameters / 32B active) , natively integrating the MoonViT multimodal vision encoder and out-of-the-box INT4 quantization, achieving a massive leap in coding performance while cutting thinking token consumption by roughly 30% compared to K2.6.Reddit Community: Where to Draw the Line Between Kimi K2.7 Code, K2.6, and K2.5The reusable value of this post is that it establishes a model-division hypothesis, rather than proving that K2.7 Code wins every real-world task: let the coding-specialized model handle repository tasks, K2.6 handle general-purpose multimodal agents, and K2.5 handle low-cost ordinary work, while using caching to control the cost of repeated context.Kimi K2.7 Code: Official GitHub Copilot Integration & Enterprise Policy SetupGitHub’s changelog documents Kimi K2.7 availability in Copilot; detail focuses on organization policy and rollout checks.Kimi K2.7 Code: Official Claude Code Integration & Multi-Tier Model MappingThe official Claude Code guide focuses on endpoint mapping, model aliases, and a controlled coding session.Kimi K2.7 Code: Official Multimodal Video Tool Calling & Agent LoopThe official multimodal example combines video input, tool calls, and a bounded agent loop; it is not a guarantee of autonomous execution.Unsiloed Benchmark: Full FastAPI Project Generation Prompt & Architectural StandardSource “Unsiloed Benchmark: Full FastAPI Project Generation Prompt & Architectural Standard” is organized as an executable task guide; its environment, inputs, and acceptance boundary follow the source.