Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaKimi K3

Artificial Analysis AA-Briefcase: Kimi K3’s Knowledge-Work Quality, Cost, and Time

Original source

Artificial Analysis

AuthorArtificial Analysis

Tabbit curation2026-08-19

Read original

Test environment

  • Benchmark: AA-Briefcase, Artificial Analysis’s private agentic knowledge-work benchmark.

  • Tasks: Complex, real-world-style input files, with deliverables including a spreadsheet, presentation, and UI mock-up.

  • Scoring: Elo synthesized from correctness, analytical quality, and presentation quality; the page does not publish the full tasks, prompts, or dataset.

  • Service: Kimi first-party API; pricing calculated at $3/$15 for Kimi K3 and $0.30/M for cached input.

Input/configuration

  • The page does not publish the complete input, system prompt, tool list, or harness version for individual tasks.

  • It records turns, output tokens, cost, and time for each task, which can be used to assess task-level cost.

Results

MetricKimi K3Comparison/interpretation
AA-Briefcase Elo1543Second, behind only Fable 5 at 1574; above GPT-5.6 Sol at 1501 and Opus 4.8 at 1347
Rubric pass rate51%Behind only Fable 5 at 56%
Analytical quality Elo1754Close to Fable 5 at 1744
Presentation quality Elo1471Below Sol at 1660 and Opus 4.8 at 1492
Average cost/task$10.57The page says this is roughly an order of magnitude higher than K2.6
Average time/task56.4 minutesAbout 2.5× Fable 5 and 3.8× Grok 4.5 high
Average output tokens120KK2.6: 42K
Average turns83Fable 5: 67; Sol: 50

Conclusion

Kimi K3 is strong at analyzing complex source material and turning it into deliverables, but it trades more turns and longer outputs for that quality. If a workflow prioritizes “correct analysis” over visual presentation, K3 is worth trying. If every task must be delivered quickly, cheaply, and with polish, routing should account for time, presentation quality, and per-task cost.

Limitations

  • AA-Briefcase is a private dataset, so readers cannot independently replay the original tasks; this document can verify the metrics and evaluation definition, but cannot claim full reproduction.

  • This is an agentic knowledge-work benchmark, not a representative test of short-form Q&A, code repair, or ordinary chat.

  • Cost and time depend on the first-party API, model version, tool calls, and task distribution; a single average cannot be used to infer every user’s bill.

  • A later snapshot in the Kimi team’s technical report lists AA-Briefcase Elo as 1548; this article’s snapshot is 1543. The date difference should be preserved rather than forcibly reconciled.

Reproduction steps

  1. Build an in-house task set containing PDFs, spreadsheets, and presentation materials, and define separate correctness, analysis, and presentation rubrics.

  2. Fix a first-party Kimi K3 API pricing snapshot, and record input/output tokens per turn, turn count, tool calls, total time, and human rework.

  3. Give the same tasks to baseline models, blind-review the deliverables, and score the three categories separately.

  4. Report means, percentiles, and cost per task; do not report only the final Elo.

Original evidence and data

The AA page publishes the benchmark’s task format and scoring dimensions, 1543 Elo, a 51% rubric pass rate, 1754 analytical Elo, 1471 presentation Elo, $10.57, 56.4 minutes, 120K output tokens, and 83 turns.

Source excerpt or observation (short compliance quote only)

The page summary says: “second only to Fable 5 … but averaging nearly an hour per task”.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Kimi K3

Use and compare models in Tabbit

Kimi K3

Related reviews

MediaGoogle / Semgrep

Kimi K3 Code Security Evaluation: Strong on the Surface, Not Precise Enough

MediaGoogle / MindStudio

Kimi K3 Real-World Coding Evaluation: Is It Really as Good as the Hype?

MediaGoogle / Simon Willison

Kimi K3 and the Pelican Benchmark: What We Can Still Learn

MediaGoogle / NxCode

Kimi K3 Benchmarks Explained: A Coding-Agent Evaluation Guide

Kimi K3

Related prompts

MediaGoogle / Business Compass LLC

Kimi K3 Prompt Engineering Guide

MediaGoogle / Together AI

Kimi K3: The Complete Developer Guide

MediaGoogle / Kimi API Platform

Kimi Prompt Best Practices

MediaGoogle / Kimi API Platform

Build an Agent with Kimi K3