Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Kimi K3 · Official source · Vendor report

Official cases show long-horizon potential and reproducibility limits

Kimi’s official release material on long-running coding, visual loops, knowledge work, and research cases.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Condition
Source: Kimi/Moonshot official technical blog.
Condition
Cases: up to 24-hour sandbox runs, 48-hour chip design, and research/knowledge-work figures.
Condition
Condition: official cases; task and harness details are incomplete.

Key data and applicable tasks

One-sentence takeaway

The official materials position Kimi K3 as a 2.8T/104B-active model suited to long-running coding, visual feedback loops, knowledge work, and research-oriented agents, while explicitly acknowledging that it still trails Claude Fable 5 and GPT-5.6 Sol overall.

Test environment

  • Model: Kimi K3, 2.8T parameters, 16/896 expert activation, 1M context, native vision.

  • Official products: Kimi.com, Kimi Work, Kimi Code, Kimi API.

  • Default settings: max thinking at release; the official footnote says evaluations used max, temperature=1.0, and top-p=1.0.

  • API pricing: $0.30/M for cache-hit input, $3/M for cache-miss input, and $15/M for output.

Input/configuration

  • Long-running agents must retain the complete thinking history; the official recommendation is to use a compatible harness such as Kimi Code and not switch models mid-session.

  • The official recommendation for overly proactive behavior is to add a more explicit system prompt or AGENTS.md boundaries.

  • In Kimi Code, select Kimi K3 with /model; the official API model name is kimi-k3.

Result data

  • GPU kernel optimization: The official materials say K3 is competitive with Fable 5 with fallback and clearly exceeds Opus 4.8, GPT-5.6 Sol, and GPT-5.5; each model was given up to 24 hours in the same sandbox to optimize, rewrite, and benchmark four tasks.

  • MiniTriton: The official case study says the model built a complete compiler from MLIR/IR to PTX codegen, reaching or exceeding Triton/torch.compile on some roofline workloads.

  • Chip design: The official materials say one 48-hour autonomous run completed design, optimization, and verification, with simulation reaching 100 MHz and 8,700+ tokens/s; these are case-study figures from the release materials.

  • Research workflow: The official materials say it reproduced I–Love–Q in about two hours, a task that normally takes 1–2 weeks, reviewing/cross-checking 20+ papers, evaluating 300+ EOS, and generating 3,000+ lines of Python plus an interactive dashboard.

  • Knowledge work: One ASIC research case used 120+ rounds of recursive improvement, 2.8k+ web searches/fetches, 1.1k+ terminal data pulls, 11k+ pages, 87 quarterly reports, and 99 original PDFs.

  • Multimodal creation: The official materials show screenshot-loop cases for games, frontend work, and CAD, as well as editing and multi-round revision from 56 source clips.

Conclusion

These cases support treating K3 as a candidate for “long-running, tool-intensive tasks that require vision or large-document context”; they do not support the conclusion that it outperforms closed models on every task. The official materials themselves place overall performance behind Fable 5 and GPT-5.6 Sol and point out that the user experience still has gaps.

Limitations

  • The page is primarily official case-study narrative and does not publish complete inputs, random seeds, tool logs, failure rates, or downloadable artifacts for each case.

  • Different benchmarks use different harnesses, including Kimi Code, Claude Code, and Codex; the official figures should not be treated as same-harness comparisons.

  • Max thinking and high token usage may cause high latency and cost; when the model is overly proactive, it may make decisions on the user's behalf under ambiguous instructions.

  • Evaluation cases are not production SLAs; enterprises should run pilots against their own repositories, data, and acceptance criteria.

Reproduction steps

  1. Select kimi-k3 in Kimi Code or the official API, and fix the harness, model version, temperature, top-p, tool permissions, and stopping conditions.

  2. Choose a real long-running task, such as a repository-level fix, screenshot-driven UI repair, or multi-document research; save the initial inputs and data snapshot.

  3. Require the complete assistant reasoning/tool-call history to be retained, and set explicit permission boundaries for overly proactive behavior.

  4. Record success rate, number of tool calls, total output tokens, elapsed time, amount of human editing, and deliverable quality; compare the same task with a known baseline model.

Original evidence and data

Verifiable fields on the official page include the model architecture, API pricing, default thinking mode, case-study runtimes/round counts/file counts, and limitations concerning thinking history, overly proactive behavior, and gaps in user experience.

Source excerpt or observation (compliance short quote only)

The official conclusion is: “While its overall performance still trails the most powerful proprietary models”.

What this supports

  • Kimi’s official release material on long-running coding, visual loops, knowledge work, research, up-to-24-hour sandbox runs, and 48-hour chip design.

What this does not support

  • Condition: official cases; task and harness details are incomplete.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Kimi official technical blog · Kimi / Moonshot AI · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Kimi K3

Compare Kimi K3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

Kimi K3 Pricing: API Costs, Subscriptions, and Budget Math

Kimi K3 pricing explained with official API rates, cache-write rules, current membership tiers, worked costs, and a practical choice framework.

Related reviews

The same K3 differed by 20 percentage points across eight harnessesA Reddit comparison of one model/provider across eight agent harnesses on 25 tasks.Kimi K3 code security evaluation: strong benchmarks do not guarantee precisionSemgrep’s IDOR code-security benchmark, separating precision, recall, and F1.Real-world coding: close on simple tasks, weaker on trap tasksCoding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.Coding-agent evidence is serious, but not “best overall”NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.Break a Kimi K3 agent loop into controlled stepsUse the Kimi API guide to connect task decomposition, tool schemas, loop control, permissions, and final checks; tools are not configured automatically in Tabbit.Turn requirements into reviewable code changes in nine stepsKimi AI’s workflow separates planning, implementation, and verification for repository changes; it does not mean the model has run your tests.Use Kimi K3 with OpenCode and Firecrawl for sourced web researchConnect Kimi K3, OpenCode, and Firecrawl MCP into a cited web-research workflow with explicit domains, permissions, and stop rules.Set up a Kimi K3 API and agent loopA sourced guide for set up a kimi k3 api and agent loop, with explicit inputs, environment, and boundaries; see the detail page for the execution path.