Semgrep’s IDOR code-security benchmark, separating precision, recall, and F1.
Google / Semgrep · Read evidenceKimi K3 · Reviews and evidence
Which Kimi K3 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
Coding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.
Google / MindStudio · Read evidenceNxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.
Google / NxCode · Read evidenceFull reviews and related reading
Selected evidence
Kimi K3 code security evaluation: strong benchmarks do not guarantee precision
Semgrep’s IDOR code-security benchmark, separating precision, recall, and F1.
- Condition
- Model/version: Kimi K3; comparison-model versions follow the source.
- Condition
- Task: IDOR detection in real open-source repositories.
- Condition
- Harness: minimal pydantic_ai agent harness; exact parameters follow the source.
Real-world coding: close on simple tasks, weaker on trap tasks
Coding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.
- Condition
- Models: Kimi K3, Kimi K2.7, and Opus 4.8.
- Condition
- Workflow: planning, implementation, and validation on issues from the same codebase.
- Condition
- Result: about 64/70 on simple tasks, just above 60 on complex tasks, and about 36% K3 failure on trap tasks.
Coding-agent evidence is serious, but not “best overall”
NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.
- Condition
- Source: NxCode; synthesis of public benchmarks.
- Condition
- Condition: mixed benchmarks, harnesses, and reasoning tiers.
- Condition
- Boundary: supports coding evaluation, not independent reproduction on one harness.
Knowledge-work quality is competitive, but cost and time per task are high
AA-Briefcase private agentic knowledge-work benchmark reporting quality, cost, time, and tokens.
- Condition
- Task: complex deliverables including spreadsheets, presentations, and UI mock-ups.
- Condition
- Snapshot: source says “last week”; exact publication date is not public.
- Condition
- Results: Elo 1543; 51% rubric pass; $10.57/task; 56.4 minutes; 120K output tokens; 83 turns.
All sources
All sources
Kimi K3 code security evaluation: strong benchmarks do not guarantee precision
Semgrep’s IDOR code-security benchmark, separating precision, recall, and F1.
- Condition
- Model/version: Kimi K3; comparison-model versions follow the source.
- Condition
- Task: IDOR detection in real open-source repositories.
- Condition
- Harness: minimal pydantic_ai agent harness; exact parameters follow the source.
Real-world coding: close on simple tasks, weaker on trap tasks
Coding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.
- Condition
- Models: Kimi K3, Kimi K2.7, and Opus 4.8.
- Condition
- Workflow: planning, implementation, and validation on issues from the same codebase.
- Condition
- Result: about 64/70 on simple tasks, just above 60 on complex tasks, and about 36% K3 failure on trap tasks.
Coding-agent evidence is serious, but not “best overall”
NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.
- Condition
- Source: NxCode; synthesis of public benchmarks.
- Condition
- Condition: mixed benchmarks, harnesses, and reasoning tiers.
- Condition
- Boundary: supports coding evaluation, not independent reproduction on one harness.
Knowledge-work quality is competitive, but cost and time per task are high
AA-Briefcase private agentic knowledge-work benchmark reporting quality, cost, time, and tokens.
- Condition
- Task: complex deliverables including spreadsheets, presentations, and UI mock-ups.
- Condition
- Snapshot: source says “last week”; exact publication date is not public.
- Condition
- Results: Elo 1543; 51% rubric pass; $10.57/task; 56.4 minutes; 120K output tokens; 83 turns.
Official cases show long-horizon potential and reproducibility limits
Kimi’s official release material on long-running coding, visual loops, knowledge work, and research cases.
- Condition
- Source: Kimi/Moonshot official technical blog.
- Condition
- Cases: up to 24-hour sandbox runs, 48-hour chip design, and research/knowledge-work figures.
- Condition
- Condition: official cases; task and harness details are incomplete.
Single SVG pelican case: useful signal, not an agent benchmark
Simon Willison used OpenRouter and llm-openrouter to generate an SVG and describe the rendered image; one personal task.
- Condition
- Model: moonshotai/kimi-k3.
- Condition
- Client: OpenRouter + llm-openrouter.
- Condition
- Task: SVG pelican riding a bicycle; 95 input tokens and 16,658 output tokens.
High capability, but slow and token-hungry: keep error bars
An X capability commentary stressing max-effort benchmarks, practical performance, and the closed-model frontier gap.
Unverified: the original source could not be rechecked.
- Condition
- Source date: page shows 2026-07-20.
- Condition
- Reasoning: usually maximum effort; high token use.
- Condition
- Evidence: personal synthesis, not a controlled rerun.
Community report: research and refusals depend on prompts and routing
A Reddit user’s personal impressions of research, coding, and refusals with Kimi K3.
Unverified: the original source could not be rechecked.
- Condition
- Source: r/LocalLLaMA; author and sample were not controlled.
- Condition
- Environment: OpenRouter/API across different use cases.
- Condition
- Content: impressions and benchmark comparisons, not a same-harness experiment.
Strong and cheap is incomplete: community disagreement on cost and capability
An LLMDevs discussion about Kimi K3’s capability, price, and practical trade-offs.
Unverified: the original source could not be rechecked.
- Condition
- Source: r/LLMDevs; discussion thread.
- Condition
- Condition: client, tasks, and pricing basis vary by user.
- Condition
- Conclusion: discussion-based opinion, not a controlled test.
Strong release benchmarks still require separate capability, cost, and deployment checks
Enter Pro’s review covering size, open weights, benchmarks, and deployment considerations.
- Condition
- Source date: page is marked 2026-07.
- Condition
- Source: third-party review; some data is cited from elsewhere.
- Condition
- Limit: not all comparisons were run under one setup.
Separate verifiable facts from unverified claims
Layer3Labs review separating model specifications, benchmark sources, and practical recommendations.
Unverified: the original source could not be rechecked.
- Condition
- Page updated: 2026-07-17.
- Condition
- Content: model specifications and public benchmark synthesis.
- Condition
- Limit: third-party article; dynamic facts need refresh.
Visual coding case study is useful, not a general success rate
Puter Developer’s visual-coding case with target/current renders and code examples.
- Condition
- Environment: Puter Developer page and browser-rendered case.
- Condition
- Task: generate or correct a visual page.
- Condition
- Evidence: case-level observation, not a large sample.
Technical report offers many numbers, but cross-harness ranking is unsafe
Kimi team technical report v2 with benchmark tables, reasoning settings, and tool conditions.
- Condition
- Version: arXiv 2607.24653v2; revised 2026-08-07.
- Condition
- Reasoning: max; temperature 1.0; top-p varies by task.
- Condition
- Tools: enabled for some visual/agent tasks.
Read top-five ranks together with their evidence labels
BenchLM places Kimi K3 benchmark rows, sources, cohorts, and evidence states together.
- Condition
- Data through 2026-08-17; 218 models; 44 displayable rows.
- Condition
- Overall 80.5/100, #5/218; top-five positions in AgenticRank, Coding, Knowledge, and Multimodal.
- Condition
- Evidence is mixed: Verified, Mixed sources, Provider exact, and other labels.
In a scoped cyber assessment it beat GLM-5.2, but ACE was 0/41
NIST/UK AISI/CAISI preliminary ExploitBench and TLO assessment.
- Condition
- ExploitBench: 41 V8/JavaScript/WebAssembly vulnerability tasks.
- Condition
- TLO: 32 steps, four subnets, about 20 hosts; 100M-token limit.
- Condition
- Results: 32%; ACE 0/41; average progress 17/32; completed 1/10 runs.
The same K3 differed by 20 percentage points across eight harnesses
A Reddit comparison of one model/provider across eight agent harnesses on 25 tasks.
- Condition
- Model: moonshotai/kimi-k3; OpenRouter; Maximum.
- Condition
- Tools: same hosted Composio MCP; eight harnesses; 25 tasks each.
- Condition
- Results: 88% highest versus 68% lowest; the source has an unresolved 25/30-task count discrepancy.
Use the comparison to design routing tests, not to declare a winner
Try Friday’s comparison of Grok 4.6 and Kimi K3 pricing, context, and deployment conditions.
- Condition
- Publication/fact-check date: 2026-08-12.
- Condition
- K3: 1M context, open weights, $3/$15, and $0.30/M cached input; Grok uses context tiers.
- Condition
- No head-to-head used the same prompt, scaffold, reasoning budget, and snapshot.
Kimi K3
Compare Kimi K3 in Tabbit
Model access, features, and permissions depend on your current client account.