Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Kimi K3 · Media / benchmark · Independent measurement

In a scoped cyber assessment it beat GLM-5.2, but ACE was 0/41

NIST/UK AISI/CAISI preliminary ExploitBench and TLO assessment.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Condition
ExploitBench: 41 V8/JavaScript/WebAssembly vulnerability tasks.
Condition
TLO: 32 steps, four subnets, about 20 hosts; 100M-token limit.
Condition
Results: 32%; ACE 0/41; average progress 17/32; completed 1/10 runs.
Condition
K3 hosted setup; complete prompt and toolchain are not public.

Key data and applicable tasks

Test environment

  • Evaluation target: Moonshot AI Kimi K3; the focus is cyber capability, not general code quality.

  • ExploitBench: 41 recent vulnerability tasks related to the V8 engine and JavaScript/WebAssembly, measuring progress from coverage/crash reproduction to arbitrary code execution (ACE).

  • The Last Ones (TLO): A simulated enterprise-network attack path with 32 steps, 4 subnets, and roughly 20 hosts; the standard limit is 100M tokens.

  • To reduce refusals, closed-source U.S. models had system-level safeguards disabled; because of its hosted setup, K3 ran only a selective cyber evaluation.

Input/configuration

  • K3’s overall cyber capability is estimated primarily from a single ExploitBench result, so its confidence interval is wider than those of other models.

  • The NIST page publishes the benchmark scale, task characteristics, and explanation of the confidence interval, but not K3’s complete system prompt, toolchain, or individual trajectories.

Results

EvaluationKimi K3Comparison/boundary
ExploitBench overall score32%GLM-5.2: 24%
ExploitBench ACE0/41Highest-severity result; the strongest models averaged 20/41
Average TLO progressStep 17/32Strongest U.S. models averaged 28.5 steps; GLM-5.2 reached step 11
TLO completion1/10 runsCompleted within the 100M-token limit

Conclusion

In this constrained evaluation, K3 clearly outperformed GLM-5.2 but still trailed the strongest cyber models, and it did not reach ACE on any of the 41 ExploitBench samples. It can be considered a security-research candidate that requires strict isolation and authorization controls, but should not be described as already having stable end-to-end attack capability.

Limitations

  • This is a preliminary assessment with a small sample, and K3 ran only a selective evaluation; the confidence interval for its overall capability estimate is wide.

  • TLO had no active defenders or alert penalties and used a preset attack path, so it differs from a real enterprise network.

  • Results with safeguards disabled for closed-source models do not represent the behavior users can directly obtain from public products.

  • These data answer questions about controlled cyber benchmarks only; they cannot be generalized to all defensive, vulnerability-audit, or production-security tasks.

Reproduction steps

  1. Obtain only the public benchmarks and version notes cited by NIST, and do so in an isolated, authorized lab environment.

  2. Fix the model version, token limit, tool permissions, and scoring script; start with low-risk vulnerability-identification and remediation tasks.

  3. Record discoveries, reproducibility, false positives, tool calls, and stopping reasons; do not extend the security evaluation to unauthorized testing of real systems.

  4. Put K3 and GLM-5.2 or other baselines on the same task set, and report confidence intervals rather than a single success rate.

Original evidence and data

The NIST page publishes 41 ExploitBench tasks, 32-step TLO, 32%/24%, ACE 0/41, 17/32, 1/10, and the 100M-token limit, and explicitly explains the differences between the benchmarks and real environments.

Source excerpt or observation (short compliance quote only)

The page’s comparison conclusion is: “Kimi K3 outperforms GLM-5.2”.

What this supports

  • NIST/UK AISI/CAISI preliminary ExploitBench and TLO assessment: 32%, ACE 0/41, average TLO progress 17/32, and 1/10 completed runs.

What this does not support

  • K3 hosted setup; complete prompt and toolchain are not public.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

NIST (in collaboration with the UK AI Security Institute) · Joint UK AISI / U.S. CAISI assessment team · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Kimi K3

Compare Kimi K3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

Kimi K3 Pricing: API Costs, Subscriptions, and Budget Math

Kimi K3 pricing explained with official API rates, cache-write rules, current membership tiers, worked costs, and a practical choice framework.

Related reviews

Kimi K3 code security evaluation: strong benchmarks do not guarantee precisionSemgrep’s IDOR code-security benchmark, separating precision, recall, and F1.Real-world coding: close on simple tasks, weaker on trap tasksCoding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.Coding-agent evidence is serious, but not “best overall”NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.Knowledge-work quality is competitive, but cost and time per task are highAA-Briefcase private agentic knowledge-work benchmark reporting quality, cost, time, and tokens.Break a Kimi K3 agent loop into controlled stepsUse the Kimi API guide to connect task decomposition, tool schemas, loop control, permissions, and final checks; tools are not configured automatically in Tabbit.Turn requirements into reviewable code changes in nine stepsKimi AI’s workflow separates planning, implementation, and verification for repository changes; it does not mean the model has run your tests.Configure Kimi K3 developer calls and multimodal inputsA sourced guide for configure kimi k3 developer calls and multimodal inputs, with explicit inputs, environment, and boundaries; see the detail page for the execution path.Choose a Kimi K3 workflow for websites, codebases, and researchA sourced guide for choose a kimi k3 workflow for websites, codebases, and research, with explicit inputs, environment, and boundaries; see the detail page for the execution path.