Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Kimi K3 · Media / benchmark · Independent measurement

Knowledge-work quality is competitive, but cost and time per task are high

AA-Briefcase private agentic knowledge-work benchmark reporting quality, cost, time, and tokens.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Condition
Task: complex deliverables including spreadsheets, presentations, and UI mock-ups.
Condition
Snapshot: source says “last week”; exact publication date is not public.
Condition
Results: Elo 1543; 51% rubric pass; $10.57/task; 56.4 minutes; 120K output tokens; 83 turns.
Condition
Service: Kimi first-party API.

Key data and applicable tasks

Test environment

  • Benchmark: AA-Briefcase, Artificial Analysis’s private agentic knowledge-work benchmark.

  • Tasks: Complex, real-world-style input files, with deliverables including a spreadsheet, presentation, and UI mock-up.

  • Scoring: Elo synthesized from correctness, analytical quality, and presentation quality; the page does not publish the full tasks, prompts, or dataset.

  • Service: Kimi first-party API; pricing calculated at $3/$15 for Kimi K3 and $0.30/M for cached input.

Input/configuration

  • The page does not publish the complete input, system prompt, tool list, or harness version for individual tasks.

  • It records turns, output tokens, cost, and time for each task, which can be used to assess task-level cost.

Results

MetricKimi K3Comparison/interpretation
AA-Briefcase Elo1543Second, behind only Fable 5 at 1574; above GPT-5.6 Sol at 1501 and Opus 4.8 at 1347
Rubric pass rate51%Behind only Fable 5 at 56%
Analytical quality Elo1754Close to Fable 5 at 1744
Presentation quality Elo1471Below Sol at 1660 and Opus 4.8 at 1492
Average cost/task$10.57The page says this is roughly an order of magnitude higher than K2.6
Average time/task56.4 minutesAbout 2.5× Fable 5 and 3.8× Grok 4.5 high
Average output tokens120KK2.6: 42K
Average turns83Fable 5: 67; Sol: 50

Conclusion

Kimi K3 is strong at analyzing complex source material and turning it into deliverables, but it trades more turns and longer outputs for that quality. If a workflow prioritizes “correct analysis” over visual presentation, K3 is worth trying. If every task must be delivered quickly, cheaply, and with polish, routing should account for time, presentation quality, and per-task cost.

Limitations

  • AA-Briefcase is a private dataset, so readers cannot independently replay the original tasks; this document can verify the metrics and evaluation definition, but cannot claim full reproduction.

  • This is an agentic knowledge-work benchmark, not a representative test of short-form Q&A, code repair, or ordinary chat.

  • Cost and time depend on the first-party API, model version, tool calls, and task distribution; a single average cannot be used to infer every user’s bill.

  • A later snapshot in the Kimi team’s technical report lists AA-Briefcase Elo as 1548; this article’s snapshot is 1543. The date difference should be preserved rather than forcibly reconciled.

Reproduction steps

  1. Build an in-house task set containing PDFs, spreadsheets, and presentation materials, and define separate correctness, analysis, and presentation rubrics.

  2. Fix a first-party Kimi K3 API pricing snapshot, and record input/output tokens per turn, turn count, tool calls, total time, and human rework.

  3. Give the same tasks to baseline models, blind-review the deliverables, and score the three categories separately.

  4. Report means, percentiles, and cost per task; do not report only the final Elo.

Original evidence and data

The AA page publishes the benchmark’s task format and scoring dimensions, 1543 Elo, a 51% rubric pass rate, 1754 analytical Elo, 1471 presentation Elo, $10.57, 56.4 minutes, 120K output tokens, and 83 turns.

Source excerpt or observation (short compliance quote only)

The page summary says: “second only to Fable 5 … but averaging nearly an hour per task”.

What this supports

  • AA-Briefcase private agentic knowledge-work benchmark reporting Elo 1543, 51% rubric pass, $10.57/task, 56.4 minutes, 120K output tokens, and 83 turns.

What this does not support

  • Does not support a public, general knowledge-work success rate because the task set is private, the source only says “last week,” and the exact publication date is undisclosed.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Kimi K3

Compare Kimi K3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

Kimi K3 Pricing: API Costs, Subscriptions, and Budget Math

Kimi K3 pricing explained with official API rates, cache-write rules, current membership tiers, worked costs, and a practical choice framework.

Related reviews

Kimi K3 code security evaluation: strong benchmarks do not guarantee precisionSemgrep’s IDOR code-security benchmark, separating precision, recall, and F1.Real-world coding: close on simple tasks, weaker on trap tasksCoding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.Coding-agent evidence is serious, but not “best overall”NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.Single SVG pelican case: useful signal, not an agent benchmarkSimon Willison used OpenRouter and llm-openrouter to generate an SVG and describe the rendered image; one personal task.Use Kimi K3 with OpenCode and Firecrawl for sourced web researchConnect Kimi K3, OpenCode, and Firecrawl MCP into a cited web-research workflow with explicit domains, permissions, and stop rules.Write Kimi API requests as testable tasksTurn the official prompting guidance into a checklist for role, context, constraints, format, and acceptance. The source does not provide one complete reusable prompt.Break a Kimi K3 agent loop into controlled stepsUse the Kimi API guide to connect task decomposition, tool schemas, loop control, permissions, and final checks; tools are not configured automatically in Tabbit.Turn requirements into reviewable code changes in nine stepsKimi AI’s workflow separates planning, implementation, and verification for repository changes; it does not mean the model has run your tests.