Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Kimi K3 · Community source · Personal experience

The same K3 differed by 20 percentage points across eight harnesses

A Reddit comparison of one model/provider across eight agent harnesses on 25 tasks.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Condition
Model: moonshotai/kimi-k3; OpenRouter; Maximum.
Condition
Tools: same hosted Composio MCP; eight harnesses; 25 tasks each.
Condition
Results: 88% highest versus 68% lowest; the source has an unresolved 25/30-task count discrepancy.
Condition
Cost uses OpenRouter list prices dated 2026-07-30.

Key data and applicable tasks

Test environment

  • Model: moonshotai/kimi-k3.

  • Provider: OpenRouter.

  • Reasoning: Maximum.

  • Tools: The same hosted Composio MCP suite, covering Gmail, Google Calendar, Google Sheets, Airtable, GitHub, Slack, Notion, Linear, and PagerDuty.

  • Tasks: 25 valid tasks per harness; the body also refers to 30 complex tasks, creating an inconsistency in scope; 200 scored runs in total.

  • Harnesses: Oh My Pi, Kimi Code, Hermes Agent, Claude Code, Pi Agent, OpenCode, Grok Build, and Codex.

Input/configuration

  • The model, provider, tools, and tasks were held constant for each harness; only the harness changed.

  • Each task was run once. Tasks covered reading and modifying business-application data, ranging from simple tasks to long workflows.

  • Costs were calculated using OpenRouter list prices from 2026-07-30; the complete task inputs, harness versions, and MCP configuration files were not published.

Results

HarnessPassedPass rateMedian completion timeTool callsCost per successful run
Oh My Pi22/2588%231.7s248$0.52
Kimi Code21/2584%281.9s301$0.64
Hermes Agent20/2580%164.0s262$0.46
Claude Code19/2576%330.5s293$1.96
Pi Agent18/2572%156.2s223$0.57
OpenCode18/2572%270.5s299$0.72
Grok Build18/2572%196.2s402$0.66
Codex17/2568%233.4s297$0.66

The highest and lowest pass rates differ by 20 percentage points. The author also observed that more tool calls do not necessarily mean better results: Grok Build made 402 calls and Pi Agent 223 calls, yet both passed 18 tasks.

Conclusion

Kimi K3’s agent results cannot be attributed to the model alone: system instructions, tool-result return, long-task context management, retries, stopping rules, and final checks all change the same model’s success rate, speed, and cost. When deploying K3, treat the harness as a measurable product component rather than comparing only model leaderboards.

Limitations

  • This was one model, one provider, and one run per task. With 25 tasks, one task changes the result by 4 percentage points, so the results do not have rigorous statistical significance.

  • Tasks were concentrated in business-application MCP workflows and cannot be generalized to coding, vision, or research tasks.

  • The author used Composio MCP, introducing tool-ecosystem and author-selection bias; the complete task list and versions were not published.

  • The same harness paired with its native model might perform differently; this does not establish that Kimi Code is always better than Claude Code/Codex.

Reproduction steps

  1. Copy the same set of authorized business-application tasks, fixing the K3 route, tool schemas, permissions, and stopping rules.

  2. Change only the harness across the eight setups, and record the version, system prompt, context compression, retry, and final-check strategies.

  3. Repeat each task at least three times, recording pass rate, p50/p95 time, tokens, tool calls, failure type, and cost per success.

  4. Report simple, medium, and difficult tasks separately so the average across 25 tasks does not hide hard-task failures.

Original evidence and data

The original Reddit post publishes the model, provider, Maximum reasoning, shared MCP tools, 25 tasks per harness, 200 runs, and the complete table of pass rates, times, tool calls, and cost per success, while explicitly listing the study’s limitations.

Source excerpt or observation (short compliance quote only)

The author reports: “The gap between the highest and lowest pass rates was 20 percentage points”.

What this supports

  • A Reddit comparison of one model/provider across eight agent harnesses on 25 tasks, reporting 88% at the high end and 68% at the low end.

What this does not support

  • Does not resolve the source’s 25-versus-30-task discrepancy or generalize to coding, vision, or research; costs follow the 2026-07-30 list prices.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit r/kimi · u/LimpComedian1317 · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Kimi K3

Compare Kimi K3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

Kimi K3 Pricing: API Costs, Subscriptions, and Budget Math

Kimi K3 pricing explained with official API rates, cache-write rules, current membership tiers, worked costs, and a practical choice framework.

Related reviews

Official cases show long-horizon potential and reproducibility limitsKimi’s official release material on long-running coding, visual loops, knowledge work, and research cases.Kimi K3 code security evaluation: strong benchmarks do not guarantee precisionSemgrep’s IDOR code-security benchmark, separating precision, recall, and F1.Real-world coding: close on simple tasks, weaker on trap tasksCoding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.Coding-agent evidence is serious, but not “best overall”NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.Break a Kimi K3 agent loop into controlled stepsUse the Kimi API guide to connect task decomposition, tool schemas, loop control, permissions, and final checks; tools are not configured automatically in Tabbit.Turn requirements into reviewable code changes in nine stepsKimi AI’s workflow separates planning, implementation, and verification for repository changes; it does not mean the model has run your tests.Use Kimi K3 with OpenCode and Firecrawl for sourced web researchConnect Kimi K3, OpenCode, and Firecrawl MCP into a cited web-research workflow with explicit domains, permissions, and stop rules.Set up a Kimi K3 API and agent loopA sourced guide for set up a kimi k3 api and agent loop, with explicit inputs, environment, and boundaries; see the detail page for the execution path.