Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityKimi K3

Reddit: Kimi K3 Compared Across Eight Agent Harnesses

Original source

Reddit r/kimi

Authoru/LimpComedian1317

Tabbit curation2026-08-19

Read original

Test environment

  • Model: moonshotai/kimi-k3.

  • Provider: OpenRouter.

  • Reasoning: Maximum.

  • Tools: The same hosted Composio MCP suite, covering Gmail, Google Calendar, Google Sheets, Airtable, GitHub, Slack, Notion, Linear, and PagerDuty.

  • Tasks: 25 valid tasks per harness; the body also refers to 30 complex tasks, creating an inconsistency in scope; 200 scored runs in total.

  • Harnesses: Oh My Pi, Kimi Code, Hermes Agent, Claude Code, Pi Agent, OpenCode, Grok Build, and Codex.

Input/configuration

  • The model, provider, tools, and tasks were held constant for each harness; only the harness changed.

  • Each task was run once. Tasks covered reading and modifying business-application data, ranging from simple tasks to long workflows.

  • Costs were calculated using OpenRouter list prices from 2026-07-30; the complete task inputs, harness versions, and MCP configuration files were not published.

Results

HarnessPassedPass rateMedian completion timeTool callsCost per successful run
Oh My Pi22/2588%231.7s248$0.52
Kimi Code21/2584%281.9s301$0.64
Hermes Agent20/2580%164.0s262$0.46
Claude Code19/2576%330.5s293$1.96
Pi Agent18/2572%156.2s223$0.57
OpenCode18/2572%270.5s299$0.72
Grok Build18/2572%196.2s402$0.66
Codex17/2568%233.4s297$0.66

The highest and lowest pass rates differ by 20 percentage points. The author also observed that more tool calls do not necessarily mean better results: Grok Build made 402 calls and Pi Agent 223 calls, yet both passed 18 tasks.

Conclusion

Kimi K3’s agent results cannot be attributed to the model alone: system instructions, tool-result return, long-task context management, retries, stopping rules, and final checks all change the same model’s success rate, speed, and cost. When deploying K3, treat the harness as a measurable product component rather than comparing only model leaderboards.

Limitations

  • This was one model, one provider, and one run per task. With 25 tasks, one task changes the result by 4 percentage points, so the results do not have rigorous statistical significance.

  • Tasks were concentrated in business-application MCP workflows and cannot be generalized to coding, vision, or research tasks.

  • The author used Composio MCP, introducing tool-ecosystem and author-selection bias; the complete task list and versions were not published.

  • The same harness paired with its native model might perform differently; this does not establish that Kimi Code is always better than Claude Code/Codex.

Reproduction steps

  1. Copy the same set of authorized business-application tasks, fixing the K3 route, tool schemas, permissions, and stopping rules.

  2. Change only the harness across the eight setups, and record the version, system prompt, context compression, retry, and final-check strategies.

  3. Repeat each task at least three times, recording pass rate, p50/p95 time, tokens, tool calls, failure type, and cost per success.

  4. Report simple, medium, and difficult tasks separately so the average across 25 tasks does not hide hard-task failures.

Original evidence and data

The original Reddit post publishes the model, provider, Maximum reasoning, shared MCP tools, 25 tasks per harness, 200 runs, and the complete table of pass rates, times, tool calls, and cost per success, while explicitly listing the study’s limitations.

Source excerpt or observation (short compliance quote only)

The author reports: “The gap between the highest and lowest pass rates was 20 percentage points”.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Kimi K3

Use and compare models in Tabbit

Kimi K3

Related reviews

MediaGoogle / Semgrep

Kimi K3 Code Security Evaluation: Strong on the Surface, Not Precise Enough

MediaGoogle / MindStudio

Kimi K3 Real-World Coding Evaluation: Is It Really as Good as the Hype?

MediaGoogle / Simon Willison

Kimi K3 and the Pelican Benchmark: What We Can Still Learn

MediaGoogle / NxCode

Kimi K3 Benchmarks Explained: A Coding-Agent Evaluation Guide

Kimi K3

Related prompts

MediaGoogle / Business Compass LLC

Kimi K3 Prompt Engineering Guide

MediaGoogle / Together AI

Kimi K3: The Complete Developer Guide

MediaGoogle / Kimi API Platform

Kimi Prompt Best Practices

MediaGoogle / Kimi API Platform

Build an Agent with Kimi K3