Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Kimi K3 · Media / benchmark · Vendor report

Technical report offers many numbers, but cross-harness ranking is unsafe

Kimi team technical report v2 with benchmark tables, reasoning settings, and tool conditions.

Media / benchmarkVendor reportEdited 2026-09-20

Test conditions

Condition
Version: arXiv 2607.24653v2; revised 2026-08-07.
Condition
Reasoning: max; temperature 1.0; top-p varies by task.
Condition
Tools: enabled for some visual/agent tasks.
Condition
Limit: model harnesses, hardware, and safeguards differ.

Key data and applicable tasks

Test environment

  • Kimi K3: 2.8T total parameters, 104B activated, native vision, 1M context.

  • Main-table baselines: Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5, and GLM-5.2.

  • Default evaluation: K3 reasoning_effort=max, temperature=1.0; top-p=0.95 for single-step knowledge/vision tasks and top-p=1.0 for agent tasks.

  • Harnesses: Kimi Code, Claude Code, and Codex; PostTrainBench is averaged over three runs on H20, ZeroBench over five, and most other vision tasks over three runs.

Input/configuration

  • BrowseComp uses 300K tokens to trigger context compaction; with the full 1M context and no management, the reported score is 90.4.

  • OfficeQA Pro renders all PDFs as images and does not provide machine-readable text.

  • MCP-Atlas uses 500 public tasks with a 100-turn limit and Gemini 3.1 Pro as the judge; AutomationBench uses 600 public tasks.

  • The report explicitly discloses that Fable 5 has fallback and GPT-5.6 Sol may trigger cyberguard, so the table is not a fully same-condition blind test.

Result data

EvaluationKimi K3How to read it
GPQA Diamond93.5Near the frontier, but GPT-5.6 Sol scores 94.1
HLE-Full (without tools/with tools)43.5 / 56.0Trails Fable/Sol on research-grade knowledge tasks
DeepSWE67.5Below Sol at 73.0 and Fable at 70.0
Terminal-Bench 2.188.3Close to Sol at 88.8
FrontierSWE81.2Second only to Fable at 86.6
ProgramBench77.8Higher in the table than Sol at 77.6 and Fable at 76.8
SWE-Marathon42.0Highest in the reported table, but task branches and harness require attention
BrowseComp91.2Higher than Sol at 90.4; see the configuration difference above
DeepSearchQA F195.0Higher than Fable at 94.2
GDPval-AA v2 Elo1686Lower than Fable at 1747 and Sol at 1736
AA-Briefcase Elo1548Lower than Fable at 1583
AutomationBench30.8Higher than Sol at 29.7 and Fable at 29.1
CharXiv (without tools/with Python)84.8 / 91.3Tools substantially change the result
Math-Vision (without tools/with Python)94.3 / 97.8Tool conditions must be distinguished
ZeroBench-main pass@5 (without tools/with Python)23.0 / 41.0Cannot be compared across tool conditions

The report also gives an internal Kimi Webdev Bench: using the same Claude Code harness as Opus 4.8, blind review produced an overall Win of 58.6%, Tie of 13.8%, and Lose of 27.6%, with Win-Lose at +31.0; this is an internal comparison, not a public benchmark.

Conclusion

The technical report supports placing K3 in the frontier-competitive range for long-horizon coding, web/tool agents, and Python-assisted vision tasks, while it still trails top closed models on knowledge-work tasks such as HLE, GDPval/AA-Briefcase, and OfficeQA. The report's greatest reuse value is that it spells out reasoning, top-p, harness, run count, and tool conditions, making a same-configuration local pilot possible.

Limitations

  • The main table is published by the model provider; internal benchmarks and cases cannot be treated as independent verification.

  • Different models use different harnesses, fallback behavior, guards, and hardware; scores cannot simply be treated as a fair competition.

  • Some third-party scores are snapshots cited from other organizations; version dates can change Elo and leaderboard standings.

  • A 1M context, tool augmentation, and agent loops can all substantially change token usage, cost, and success rate.

Reproduction steps

  1. Fix the Kimi K3 version, max, and temperature 1.0, and set top-p to 0.95/1.0 by task type.

  2. Record the Kimi Code/Claude Code/Codex harness, tool permissions, context-compaction threshold, and hardware.

  3. Run public benchmarks at least three times according to the official version; for vision tasks record whether Python is enabled, and run ZeroBench five times.

  4. Store each benchmark's inputs, stopping rules, failures/refusals, tokens, cost, and results separately; do not merge figures from different harnesses into a single conclusion.

Original evidence and data

The technical report publicly provides the main table, evaluation configuration, third-party citation notes, internal benchmark definitions, tool conditions, and result explanations. These fields are sufficient to verify the “comparability boundaries caused by configuration,” but insufficient to reproduce every private internal task.

Source excerpt or observation (compliance short quote only)

The report's table header explicitly says: “All maxed out on thinking effort: max or xhigh.”

What this supports

  • Kimi team technical report v2 (arXiv 2607.24653v2, revised 2026-08-07) with benchmark tables, max reasoning, temperature 1.0, and tool conditions.

What this does not support

  • Limit: model harnesses, hardware, and safeguards differ.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

arXiv · Kimi Team (more than 302 authors are not all expanded on the abstract page) · Original publication date 2026-07-27 · Site edit date 2026-09-20

Open original source

Kimi K3

Compare Kimi K3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

Kimi K3 Pricing: API Costs, Subscriptions, and Budget Math

Kimi K3 pricing explained with official API rates, cache-write rules, current membership tiers, worked costs, and a practical choice framework.

Related reviews

Strong release benchmarks still require separate capability, cost, and deployment checksEnter Pro’s review covering size, open weights, benchmarks, and deployment considerations.Separate verifiable facts from unverified claimsLayer3Labs review separating model specifications, benchmark sources, and practical recommendations.Read top-five ranks together with their evidence labelsBenchLM places Kimi K3 benchmark rows, sources, cohorts, and evidence states together.High capability, but slow and token-hungry: keep error barsAn X capability commentary stressing max-effort benchmarks, practical performance, and the closed-model frontier gap.Write Kimi API requests as testable tasksTurn the official prompting guidance into a checklist for role, context, constraints, format, and acceptance. The source does not provide one complete reusable prompt.Turn Kimi K3 prompting advice into executable constraintsA sourced guide for turn kimi k3 prompting advice into executable constraints, with explicit inputs, environment, and boundaries; see the detail page for the execution path.Use Kimi K3 with OpenCode and Firecrawl for sourced web researchConnect Kimi K3, OpenCode, and Firecrawl MCP into a cited web-research workflow with explicit domains, permissions, and stop rules.Break a Kimi K3 agent loop into controlled stepsUse the Kimi API guide to connect task decomposition, tool schemas, loop control, permissions, and final checks; tools are not configured automatically in Tabbit.