Kimi K3

Kimi K3 · Reviews and evidence

Which Kimi K3 conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

Full reviews and related reading

Read the full analysis

Pricing · English

Kimi K3 Pricing: API Costs, Subscriptions, and Budget Math

Kimi K3 pricing explained with official API rates, cache-write rules, current membership tiers, worked costs, and a practical choice framework.

Selected evidence

Media / benchmarkIndependent measurement

Kimi K3 code security evaluation: strong benchmarks do not guarantee precision

Semgrep’s IDOR code-security benchmark, separating precision, recall, and F1.

SourceGoogle / Semgrep
PublishedUnknown
Collected2026-08-18
Condition
Model/version: Kimi K3; comparison-model versions follow the source.
Condition
Task: IDOR detection in real open-source repositories.
Condition
Harness: minimal pydantic_ai agent harness; exact parameters follow the source.
Coding
Media / benchmarkIndependent measurement

Real-world coding: close on simple tasks, weaker on trap tasks

Coding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.

SourceGoogle / MindStudio
PublishedUnknown
Collected2026-08-18
Condition
Models: Kimi K3, Kimi K2.7, and Opus 4.8.
Condition
Workflow: planning, implementation, and validation on issues from the same codebase.
Condition
Result: about 64/70 on simple tasks, just above 60 on complex tasks, and about 36% K3 failure on trap tasks.
Coding
Media / benchmarkEditorial analysis

Coding-agent evidence is serious, but not “best overall”

NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.

SourceGoogle / NxCode
PublishedUnknown
Collected2026-08-18
Condition
Source: NxCode; synthesis of public benchmarks.
Condition
Condition: mixed benchmarks, harnesses, and reasoning tiers.
Condition
Boundary: supports coding evaluation, not independent reproduction on one harness.
Coding
Media / benchmarkIndependent measurement

Knowledge-work quality is competitive, but cost and time per task are high

AA-Briefcase private agentic knowledge-work benchmark reporting quality, cost, time, and tokens.

SourceArtificial Analysis
PublishedUnknown
Collected2026-08-18
Condition
Task: complex deliverables including spreadsheets, presentations, and UI mock-ups.
Condition
Snapshot: source says “last week”; exact publication date is not public.
Condition
Results: Elo 1543; 51% rubric pass; $10.57/task; 56.4 minutes; 120K output tokens; 83 turns.
Information extraction

All sources

All sources

17 / 17
Media / benchmarkIndependent measurement

Kimi K3 code security evaluation: strong benchmarks do not guarantee precision

Semgrep’s IDOR code-security benchmark, separating precision, recall, and F1.

SourceGoogle / Semgrep
PublishedUnknown
Collected2026-08-18
Condition
Model/version: Kimi K3; comparison-model versions follow the source.
Condition
Task: IDOR detection in real open-source repositories.
Condition
Harness: minimal pydantic_ai agent harness; exact parameters follow the source.
Coding
Media / benchmarkIndependent measurement

Real-world coding: close on simple tasks, weaker on trap tasks

Coding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.

SourceGoogle / MindStudio
PublishedUnknown
Collected2026-08-18
Condition
Models: Kimi K3, Kimi K2.7, and Opus 4.8.
Condition
Workflow: planning, implementation, and validation on issues from the same codebase.
Condition
Result: about 64/70 on simple tasks, just above 60 on complex tasks, and about 36% K3 failure on trap tasks.
Coding
Media / benchmarkEditorial analysis

Coding-agent evidence is serious, but not “best overall”

NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.

SourceGoogle / NxCode
PublishedUnknown
Collected2026-08-18
Condition
Source: NxCode; synthesis of public benchmarks.
Condition
Condition: mixed benchmarks, harnesses, and reasoning tiers.
Condition
Boundary: supports coding evaluation, not independent reproduction on one harness.
Coding
Media / benchmarkIndependent measurement

Knowledge-work quality is competitive, but cost and time per task are high

AA-Briefcase private agentic knowledge-work benchmark reporting quality, cost, time, and tokens.

SourceArtificial Analysis
PublishedUnknown
Collected2026-08-18
Condition
Task: complex deliverables including spreadsheets, presentations, and UI mock-ups.
Condition
Snapshot: source says “last week”; exact publication date is not public.
Condition
Results: Elo 1543; 51% rubric pass; $10.57/task; 56.4 minutes; 120K output tokens; 83 turns.
Information extraction
OfficialVendor report

Official cases show long-horizon potential and reproducibility limits

Kimi’s official release material on long-running coding, visual loops, knowledge work, and research cases.

SourceKimi official technical blog
PublishedUnknown
Collected2026-08-18
Condition
Source: Kimi/Moonshot official technical blog.
Condition
Cases: up to 24-hour sandbox runs, 48-hour chip design, and research/knowledge-work figures.
Condition
Condition: official cases; task and harness details are incomplete.
Agent
Media / benchmarkPersonal experience

Single SVG pelican case: useful signal, not an agent benchmark

Simon Willison used OpenRouter and llm-openrouter to generate an SVG and describe the rendered image; one personal task.

SourceGoogle / Simon Willison
PublishedUnknown
Collected2026-08-18
Condition
Model: moonshotai/kimi-k3.
Condition
Client: OpenRouter + llm-openrouter.
Condition
Task: SVG pelican riding a bicycle; 95 input tokens and 16,658 output tokens.
Visual generation
CommunityPersonal experience

High capability, but slow and token-hungry: keep error bars

An X capability commentary stressing max-effort benchmarks, practical performance, and the closed-model frontier gap.

SourceX
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Condition
Source date: page shows 2026-07-20.
Condition
Reasoning: usually maximum effort; high token use.
Condition
Evidence: personal synthesis, not a controlled rerun.
Reasoning
CommunityPersonal experience

Community report: research and refusals depend on prompts and routing

A Reddit user’s personal impressions of research, coding, and refusals with Kimi K3.

SourceReddit / r/LocalLLaMA
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Condition
Source: r/LocalLLaMA; author and sample were not controlled.
Condition
Environment: OpenRouter/API across different use cases.
Condition
Content: impressions and benchmark comparisons, not a same-harness experiment.
Reasoning
CommunityPersonal experience

Strong and cheap is incomplete: community disagreement on cost and capability

An LLMDevs discussion about Kimi K3’s capability, price, and practical trade-offs.

SourceReddit / r/LLMDevs
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Condition
Source: r/LLMDevs; discussion thread.
Condition
Condition: client, tasks, and pricing basis vary by user.
Condition
Conclusion: discussion-based opinion, not a controlled test.
Reasoning
Media / benchmarkEditorial analysis

Strong release benchmarks still require separate capability, cost, and deployment checks

Enter Pro’s review covering size, open weights, benchmarks, and deployment considerations.

SourceGoogle / Enter Pro
PublishedUnknown
Collected2026-08-18
Condition
Source date: page is marked 2026-07.
Condition
Source: third-party review; some data is cited from elsewhere.
Condition
Limit: not all comparisons were run under one setup.
Reasoning
Media / benchmarkEditorial analysis

Separate verifiable facts from unverified claims

Layer3Labs review separating model specifications, benchmark sources, and practical recommendations.

SourceGoogle / Layer3Labs
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Condition
Page updated: 2026-07-17.
Condition
Content: model specifications and public benchmark synthesis.
Condition
Limit: third-party article; dynamic facts need refresh.
Reasoning
Media / benchmarkIndependent measurement

Visual coding case study is useful, not a general success rate

Puter Developer’s visual-coding case with target/current renders and code examples.

SourceGoogle / Puter Developer
PublishedUnknown
Collected2026-08-18
Condition
Environment: Puter Developer page and browser-rendered case.
Condition
Task: generate or correct a visual page.
Condition
Evidence: case-level observation, not a large sample.
Visual generation
Media / benchmarkVendor report

Technical report offers many numbers, but cross-harness ranking is unsafe

Kimi team technical report v2 with benchmark tables, reasoning settings, and tool conditions.

SourcearXiv
Published2026-07-27
Collected2026-08-18
Condition
Version: arXiv 2607.24653v2; revised 2026-08-07.
Condition
Reasoning: max; temperature 1.0; top-p varies by task.
Condition
Tools: enabled for some visual/agent tasks.
Reasoning
Media / benchmarkEditorial analysis

Read top-five ranks together with their evidence labels

BenchLM places Kimi K3 benchmark rows, sources, cohorts, and evidence states together.

SourceBenchLM
Published2026-08-17
Collected2026-08-18
Condition
Data through 2026-08-17; 218 models; 44 displayable rows.
Condition
Overall 80.5/100, #5/218; top-five positions in AgenticRank, Coding, Knowledge, and Multimodal.
Condition
Evidence is mixed: Verified, Mixed sources, Provider exact, and other labels.
Reasoning
Media / benchmarkIndependent measurement

In a scoped cyber assessment it beat GLM-5.2, but ACE was 0/41

NIST/UK AISI/CAISI preliminary ExploitBench and TLO assessment.

SourceNIST (in collaboration with the UK AI Security Institute)
PublishedUnknown
Collected2026-08-18
Condition
ExploitBench: 41 V8/JavaScript/WebAssembly vulnerability tasks.
Condition
TLO: 32 steps, four subnets, about 20 hosts; 100M-token limit.
Condition
Results: 32%; ACE 0/41; average progress 17/32; completed 1/10 runs.
Coding
CommunityPersonal experience

The same K3 differed by 20 percentage points across eight harnesses

A Reddit comparison of one model/provider across eight agent harnesses on 25 tasks.

SourceReddit r/kimi
PublishedUnknown
Collected2026-08-18
Condition
Model: moonshotai/kimi-k3; OpenRouter; Maximum.
Condition
Tools: same hosted Composio MCP; eight harnesses; 25 tasks each.
Condition
Results: 88% highest versus 68% lowest; the source has an unresolved 25/30-task count discrepancy.
Agent
Media / benchmarkEditorial analysis

Use the comparison to design routing tests, not to declare a winner

Try Friday’s comparison of Grok 4.6 and Kimi K3 pricing, context, and deployment conditions.

SourceTry Friday AI Research
Published2026-08-12
Collected2026-08-18
Condition
Publication/fact-check date: 2026-08-12.
Condition
K3: 1M context, open weights, $3/$15, and $0.30/M cached input; Grok uses context tiers.
Condition
No head-to-head used the same prompt, scaffold, reasoning budget, and snapshot.
Cost

Kimi K3

Compare Kimi K3 in Tabbit

Model access, features, and permissions depend on your current client account.