BenchLM's dynamic ledger shows publicly sourced results for Kimi K2.5 across Coding, Agentic, Reasoning, Multimodal, and other benchmark categories, but its total score and ranking use a custom aggregation; verification should compare individual original benchmarks, tool modes, and context strategies.
Page status: Data through 2026-08-17; the page shows 21 source-displayable rows (the dynamic directory will update).
Evidence: A mix of Benchmark exact, Provider exact, and Secondary exact; the page includes SWE-Bench, Terminal-Bench, BrowseComp, MMMU, GPQA, and others.
Aggregation: The page's custom category scores/weights and overall ranking are not the result of one unified run.
Entries come from the Kimi model card, Fireworks, and SWE/Terminal/AA and other leaderboards; Thinking, tools, context, and repeat counts differ across benchmarks. When using the ledger, open the line-level sources first and fix the K2.5 snapshot, mode, parser, and harness.
Representative results visible on the page include: SWE-Bench Verified 76.8%, SWE-Bench Pro 50.7%, SWE Multilingual 73.0%, Terminal-Bench 2.0 50.8%, BrowseComp 60.6%, BrowseComp (context management) 74.9%, Agent Swarm 78.4%, MMMU-Pro 78.5%, LongBench v2 61.0%, and AA-LCR 70.0%.
K2.5's task strengths are concentrated in multimodality, Agentic Search, visual documents, and parallelizable coding. Whether context management, Thinking/non-thinking, and Swarm are enabled can substantially change scores, so tasks must be re-run for the target product.
BenchLM's total score and ranking are affected by custom weights, directory updates, and mixed sources and cannot replace original metrics.
Kimi's official figures use internal prompts/harnesses and an independent Swarm configuration; results from different sources are not necessarily directly comparable.
Some results come from providers or secondary sources; the evidence level must be tracked for each row.
Select the target task first: coding, vision, search, office work, long context, or Swarm.
Lock the model mode, temperature/top_p, tools, context management, maximum steps, and repeat count for each item.
Save the original prompt, input media, tool trace, output, score, cost, and failure samples.
Re-run with the same harness against at least one comparison model, using the BenchLM page only as an index and cross-check.
Kimi K2.5