Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaKimi K3

BenchLM: Kimi K3’s Publicly Verifiable Benchmark Ledger

Original source

BenchLM

AuthorBenchLM

Source date2026-08-17

Tabbit curation2026-08-19

Read original

Test environment

  • BenchLM model profile and benchmark ledger; the page places each row’s score, comparison, weight, cohort, and evidence status together.

  • The page’s data is current through 2026-08-17 and covers a ranking view of 218 models.

  • Pricing: $3/M input, $15/M output, and $0.30/M cached input.

Input/configuration

  • The page shows 44 benchmark rows with displayable sources; evidence statuses include Verified, Mixed sources, Provider exact, Benchmark exact, and Secondary exact.

  • Many of K3’s coding/agentic figures come from Kimi’s official published tables. BenchLM is an evidence ledger and cross-model comparison layer, not a rerun using a unified harness.

Results

CategoryKimi K3Ranking/evidence
Overall score80.5/100#5/218; 44 source-displayable rows
AgenticRank73.9#5/131, 10 benchmarks, Verified
CodingRank78.0#5/135, 13 benchmarks, Mixed sources
Knowledge84.8#5/57, 4 benchmarks, Verified
Multimodal87.9#4/33, 13 benchmarks, Verified
Reasoning/Math/Multilingual/InstructionNot measuredNo usable benchmark rows on the page

Key ledger rows include: DeepSWE 67.5 (below GPT-5.6 Sol at 72.7), Terminal-Bench 2.1 88.3 (below Sol at 91.9), BrowseComp 91.2 (below Sol at 92.2), AutomationBench 30.8 (below GLM-5.3 at 48.2), APEX-Agents 37.6 (below Grok 4.6 at 57.5), MathVision 94.3/97.8 (without tools/with Python), and ExploitBench 32 (with a source chain pointing to NIST/UK AISI/CAISI).

Conclusion

BenchLM’s value is not announcing another “first”, but breaking K3’s ranking down into traceable sources: it is in the top five for Agentic, Coding, Knowledge, and Multimodal. However, mixing Provider exact, Benchmark exact, and Verified figures can easily lead readers to misinterpret all of them as independently rerun results.

Limitations

  • The page does not provide K3’s unified inputs, complete harness, run logs, or raw outputs; it cannot replace an independent controlled evaluation.

  • Many of the 44 rows are provider exact, and some categories are entirely unmeasured; the overall score does not cover all capabilities.

  • Rankings and Elo/source records change as new models, sources, and benchmark versions are added or updated; the collection date should be retained.

Reproduction steps

  1. Save a page snapshot and each row’s source link, evidence status, and date.

  2. Use only Verified or Benchmark exact rows as candidates for review, opening the original leaderboard/report for each one.

  3. Build a local held-out set with the same categories for the actual selection task, and report K3 and baseline results under the same harness.

  4. Do not write Provider exact figures as independent experimental results; retain the original evidence labels.

Original evidence and data

The page publishes the overall score of 80.5, #5/218, 44 source-displayable rows, category scores/rankings, pricing, and the comparison and evidence type for each benchmark; it links to Kimi’s official blog, NIST/UK AISI/CAISI, and individual benchmark pages.

Source excerpt or observation (short compliance quote only)

The page’s evidence summary is: “Evidence: Supported; this profile shows 44 source-displayable benchmark rows.”

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Kimi K3

Use and compare models in Tabbit

Kimi K3

Related reviews

MediaGoogle / Semgrep

Kimi K3 Code Security Evaluation: Strong on the Surface, Not Precise Enough

MediaGoogle / MindStudio

Kimi K3 Real-World Coding Evaluation: Is It Really as Good as the Hype?

MediaGoogle / Simon Willison

Kimi K3 and the Pelican Benchmark: What We Can Still Learn

MediaGoogle / NxCode

Kimi K3 Benchmarks Explained: A Coding-Agent Evaluation Guide

Kimi K3

Related prompts

MediaGoogle / Business Compass LLC

Kimi K3 Prompt Engineering Guide

MediaGoogle / Together AI

Kimi K3: The Complete Developer Guide

MediaGoogle / Kimi API Platform

Kimi Prompt Best Practices

MediaGoogle / Kimi API Platform

Build an Agent with Kimi K3