Kimi K2.5 · Media / benchmark · Independent measurement
BenchLM's Kimi K2.5 ledger, current through 2026-08-17, aggregates Coding, Agentic, Reasoning, and Multimodal sources, but its total score and ranking are custom aggregates.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
BenchLM's dynamic ledger shows publicly sourced results for Kimi K2.5 across Coding, Agentic, Reasoning, Multimodal, and other benchmark categories, but its total score and ranking use a custom aggregation; verification should compare individual original benchmarks, tool modes, and context strategies.
Page status: Data through 2026-08-17; the page shows 21 source-displayable rows (the dynamic directory will update).
Evidence: A mix of Benchmark exact, Provider exact, and Secondary exact; the page includes SWE-Bench, Terminal-Bench, BrowseComp, MMMU, GPQA, and others.
Aggregation: The page's custom category scores/weights and overall ranking are not the result of one unified run.
Entries come from the Kimi model card, Fireworks, and SWE/Terminal/AA and other leaderboards; Thinking, tools, context, and repeat counts differ across benchmarks. When using the ledger, open the line-level sources first and fix the K2.5 snapshot, mode, parser, and harness.
Representative results visible on the page include: SWE-Bench Verified 76.8%, SWE-Bench Pro 50.7%, SWE Multilingual 73.0%, Terminal-Bench 2.0 50.8%, BrowseComp 60.6%, BrowseComp (context management) 74.9%, Agent Swarm 78.4%, MMMU-Pro 78.5%, LongBench v2 61.0%, and AA-LCR 70.0%.
K2.5's task strengths are concentrated in multimodality, Agentic Search, visual documents, and parallelizable coding. Whether context management, Thinking/non-thinking, and Swarm are enabled can substantially change scores, so tasks must be re-run for the target product.
BenchLM's total score and ranking are affected by custom weights, directory updates, and mixed sources and cannot replace original metrics.
Kimi's official figures use internal prompts/harnesses and an independent Swarm configuration; results from different sources are not necessarily directly comparable.
Some results come from providers or secondary sources; the evidence level must be tracked for each row.
Select the target task first: coding, vision, search, office work, long context, or Swarm.
Lock the model mode, temperature/top_p, tools, context management, maximum steps, and repeat count for each item.
Save the original prompt, input media, tool trace, output, score, cost, and failure samples.
Re-run with the same harness against at least one comparison model, using the BenchLM page only as an index and cross-check.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
BenchLM · BenchLM · Original publication date 2026-08-17 · Site edit date 2026-09-20
Open original sourceKimi K2.5
Download the Tabbit client to check model access