Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaKimi K3

Kimi K3 Technical Report: Complete Benchmark Table and Evaluation Configuration

Original source

arXiv

AuthorKimi Team (more than 302 authors are not all expanded on the abstract page)

Source date2026-07-27

Tabbit curation2026-08-19

Read original

Test environment

  • Kimi K3: 2.8T total parameters, 104B activated, native vision, 1M context.

  • Main-table baselines: Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5, and GLM-5.2.

  • Default evaluation: K3 reasoning_effort=max, temperature=1.0; top-p=0.95 for single-step knowledge/vision tasks and top-p=1.0 for agent tasks.

  • Harnesses: Kimi Code, Claude Code, and Codex; PostTrainBench is averaged over three runs on H20, ZeroBench over five, and most other vision tasks over three runs.

Input/configuration

  • BrowseComp uses 300K tokens to trigger context compaction; with the full 1M context and no management, the reported score is 90.4.

  • OfficeQA Pro renders all PDFs as images and does not provide machine-readable text.

  • MCP-Atlas uses 500 public tasks with a 100-turn limit and Gemini 3.1 Pro as the judge; AutomationBench uses 600 public tasks.

  • The report explicitly discloses that Fable 5 has fallback and GPT-5.6 Sol may trigger cyberguard, so the table is not a fully same-condition blind test.

Result data

EvaluationKimi K3How to read it
GPQA Diamond93.5Near the frontier, but GPT-5.6 Sol scores 94.1
HLE-Full (without tools/with tools)43.5 / 56.0Trails Fable/Sol on research-grade knowledge tasks
DeepSWE67.5Below Sol at 73.0 and Fable at 70.0
Terminal-Bench 2.188.3Close to Sol at 88.8
FrontierSWE81.2Second only to Fable at 86.6
ProgramBench77.8Higher in the table than Sol at 77.6 and Fable at 76.8
SWE-Marathon42.0Highest in the reported table, but task branches and harness require attention
BrowseComp91.2Higher than Sol at 90.4; see the configuration difference above
DeepSearchQA F195.0Higher than Fable at 94.2
GDPval-AA v2 Elo1686Lower than Fable at 1747 and Sol at 1736
AA-Briefcase Elo1548Lower than Fable at 1583
AutomationBench30.8Higher than Sol at 29.7 and Fable at 29.1
CharXiv (without tools/with Python)84.8 / 91.3Tools substantially change the result
Math-Vision (without tools/with Python)94.3 / 97.8Tool conditions must be distinguished
ZeroBench-main pass@5 (without tools/with Python)23.0 / 41.0Cannot be compared across tool conditions

The report also gives an internal Kimi Webdev Bench: using the same Claude Code harness as Opus 4.8, blind review produced an overall Win of 58.6%, Tie of 13.8%, and Lose of 27.6%, with Win-Lose at +31.0; this is an internal comparison, not a public benchmark.

Conclusion

The technical report supports placing K3 in the frontier-competitive range for long-horizon coding, web/tool agents, and Python-assisted vision tasks, while it still trails top closed models on knowledge-work tasks such as HLE, GDPval/AA-Briefcase, and OfficeQA. The report's greatest reuse value is that it spells out reasoning, top-p, harness, run count, and tool conditions, making a same-configuration local pilot possible.

Limitations

  • The main table is published by the model provider; internal benchmarks and cases cannot be treated as independent verification.

  • Different models use different harnesses, fallback behavior, guards, and hardware; scores cannot simply be treated as a fair competition.

  • Some third-party scores are snapshots cited from other organizations; version dates can change Elo and leaderboard standings.

  • A 1M context, tool augmentation, and agent loops can all substantially change token usage, cost, and success rate.

Reproduction steps

  1. Fix the Kimi K3 version, max, and temperature 1.0, and set top-p to 0.95/1.0 by task type.

  2. Record the Kimi Code/Claude Code/Codex harness, tool permissions, context-compaction threshold, and hardware.

  3. Run public benchmarks at least three times according to the official version; for vision tasks record whether Python is enabled, and run ZeroBench five times.

  4. Store each benchmark's inputs, stopping rules, failures/refusals, tokens, cost, and results separately; do not merge figures from different harnesses into a single conclusion.

Original evidence and data

The technical report publicly provides the main table, evaluation configuration, third-party citation notes, internal benchmark definitions, tool conditions, and result explanations. These fields are sufficient to verify the “comparability boundaries caused by configuration,” but insufficient to reproduce every private internal task.

Source excerpt or observation (compliance short quote only)

The report's table header explicitly says: “All maxed out on thinking effort: max or xhigh.”

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Kimi K3

Use and compare models in Tabbit

Kimi K3

Related reviews

MediaGoogle / Semgrep

Kimi K3 Code Security Evaluation: Strong on the Surface, Not Precise Enough

MediaGoogle / MindStudio

Kimi K3 Real-World Coding Evaluation: Is It Really as Good as the Hype?

MediaGoogle / Simon Willison

Kimi K3 and the Pelican Benchmark: What We Can Still Learn

MediaGoogle / NxCode

Kimi K3 Benchmarks Explained: A Coding-Agent Evaluation Guide

Kimi K3

Related prompts

MediaGoogle / Business Compass LLC

Kimi K3 Prompt Engineering Guide

MediaGoogle / Together AI

Kimi K3: The Complete Developer Guide

MediaGoogle / Kimi API Platform

Kimi Prompt Best Practices

MediaGoogle / Kimi API Platform

Build an Agent with Kimi K3