Model: moonshotai/kimi-k3.
Provider: OpenRouter.
Reasoning: Maximum.
Tools: The same hosted Composio MCP suite, covering Gmail, Google Calendar, Google Sheets, Airtable, GitHub, Slack, Notion, Linear, and PagerDuty.
Tasks: 25 valid tasks per harness; the body also refers to 30 complex tasks, creating an inconsistency in scope; 200 scored runs in total.
Harnesses: Oh My Pi, Kimi Code, Hermes Agent, Claude Code, Pi Agent, OpenCode, Grok Build, and Codex.
The model, provider, tools, and tasks were held constant for each harness; only the harness changed.
Each task was run once. Tasks covered reading and modifying business-application data, ranging from simple tasks to long workflows.
Costs were calculated using OpenRouter list prices from 2026-07-30; the complete task inputs, harness versions, and MCP configuration files were not published.
| Harness | Passed | Pass rate | Median completion time | Tool calls | Cost per successful run |
|---|---|---|---|---|---|
| Oh My Pi | 22/25 | 88% | 231.7s | 248 | $0.52 |
| Kimi Code | 21/25 | 84% | 281.9s | 301 | $0.64 |
| Hermes Agent | 20/25 | 80% | 164.0s | 262 | $0.46 |
| Claude Code | 19/25 | 76% | 330.5s | 293 | $1.96 |
| Pi Agent | 18/25 | 72% | 156.2s | 223 | $0.57 |
| OpenCode | 18/25 | 72% | 270.5s | 299 | $0.72 |
| Grok Build | 18/25 | 72% | 196.2s | 402 | $0.66 |
| Codex | 17/25 | 68% | 233.4s | 297 | $0.66 |
The highest and lowest pass rates differ by 20 percentage points. The author also observed that more tool calls do not necessarily mean better results: Grok Build made 402 calls and Pi Agent 223 calls, yet both passed 18 tasks.
Kimi K3’s agent results cannot be attributed to the model alone: system instructions, tool-result return, long-task context management, retries, stopping rules, and final checks all change the same model’s success rate, speed, and cost. When deploying K3, treat the harness as a measurable product component rather than comparing only model leaderboards.
This was one model, one provider, and one run per task. With 25 tasks, one task changes the result by 4 percentage points, so the results do not have rigorous statistical significance.
Tasks were concentrated in business-application MCP workflows and cannot be generalized to coding, vision, or research tasks.
The author used Composio MCP, introducing tool-ecosystem and author-selection bias; the complete task list and versions were not published.
The same harness paired with its native model might perform differently; this does not establish that Kimi Code is always better than Claude Code/Codex.
Copy the same set of authorized business-application tasks, fixing the K3 route, tool schemas, permissions, and stopping rules.
Change only the harness across the eight setups, and record the version, system prompt, context compression, retry, and final-check strategies.
Repeat each task at least three times, recording pass rate, p50/p95 time, tokens, tool calls, failure type, and cost per success.
Report simple, medium, and difficult tasks separately so the average across 25 tasks does not hide hard-task failures.
The original Reddit post publishes the model, provider, Maximum reasoning, shared MCP tools, 25 tasks per harness, 200 runs, and the complete table of pass rates, times, tool calls, and cost per success, while explicitly listing the study’s limitations.
The author reports: “The gap between the highest and lowest pass rates was 20 percentage points”.
Kimi K3