Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Kimi K2.5 · Media / benchmark · Independent measurement

Fireworks' Quality Comparison of the Official Kimi K2.5 API and Deployment Stack

Fireworks reruns Kimi through the official API and shows that chat templates, EOS, reasoning_content, sampling, and load errors can change tool-call quality.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
Fireworks quality report comparing the Kimi official API with the Fireworks deployment stack; exact date is unpublished.
Published conditions
It observes effects from chat templates, EOS/thinking boundaries, null reasoning_content, sampling, and load errors on tool calls.

Key data and applicable tasks

One-sentence takeaway

Using the official Kimi API as a reference, Fireworks found that production quality is determined by more than the model: K2.5's chat template, EOS/thinking boundary, null reasoning_content, sampling parameters, and load errors can all change tool-call results, while Fireworks' representative scores are close to the official figures with minor differences.

Test environment

  • Comparison: Kimi model card, Kimi official API measured by Fireworks, and Fireworks API.

  • Quality layer: prompt formatting, stream/non-stream tool calling, grammar-constrained generation, numerical correctness, timeout/disconnect/error codes, and SDK/proxy compatibility.

  • Evaluation layer: deterministic unit tests, one-turn, agentic multi-turn, multimodal, and periodic regression; KVV, SWE-Agent, Terminus-2, and Eval Protocol.

  • Repeats: Most results avg@3; Tau-Bench avg@4; SWE-Bench avg@2.

Input/configuration

The Fireworks report states that SWE-Bench uses SWE-Agent, Terminal-Bench 2 uses Terminus 2, and AIME/MMMU/OCR/K2VV use Kimi KVV. For SWE-Bench, Fireworks raised max output to 32k, while the official harness does not specify max_tokens; Terminal-Bench is forced to non-thinking mode.

Results

EvaluationKimi Model CardOfficial API (measured by Fireworks)Fireworks API
SWE-Bench76.876.876.8
Terminal-Bench 2 (non-thinking)50.8Not publicly disclosed50.6
Tau2 AirlineNot publicly disclosed6568
AIME 202596.195.795.0
K2VV ToolCall F1 (non-thinking)84.0Not publicly disclosed83.9
MMMU Pro Vision77.4Not publicly disclosed77.9
OCRBench91.0Not publicly disclosed91.7

Conclusion

Fireworks' measurements support the view that “correct deployment can approach official results,” but they also show that production systems must handle thinking-phase EOS, null reasoning_content, sampling parameters, empty 200 responses, and multi-turn tool calls. Product selection should evaluate the complete serving stack rather than compare model cards alone.

Limitations

  • Fireworks has a provider perspective; code/fix details and some official KVV results depend in part on its own service.

  • Most scores use only a small number of repeats (2–4) and cannot represent variance across all tasks.

  • Promotional comparisons such as “1/10 the cost” and “2–3× the speed” depend on provider conditions and are not universal guarantees of price or throughput.

  • Terminal non-thinking, the SWE custom harness, and official API conditions are not identical; cross-column comparisons require caution.

Reproduction steps

  1. For the same prompt, model snapshot, tool schema, sampling settings, and max tokens, call the official API, Fireworks, and a self-hosted endpoint separately.

  2. Run single-turn/multi-turn, stream/non-stream, Thinking/Instant, vision, and grammar-constrained tool tests.

  3. Verify the chat template, reasoning_content null filtering, EOS guard, HTTP 429/empty 200 responses, timeouts, and retries.

  4. Report score, call success rate, error type, latency, and cost for each harness; avoid summarizing deployment quality with a single SWE score.

What this supports

  • It supports including provider and harness configuration in K2.5 tests

What this does not support

  • It supports including provider and harness configuration in K2.5 tests, not representing every cloud or local deployment.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Fireworks AI · Fireworks AI · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Kimi K2.5

Compare Kimi K2.5 in Tabbit

Download the Tabbit client to check model access

Related reviews

BenchLM's Public Benchmark Ledger and Task Stratification for Kimi K2.5BenchLM's Kimi K2.5 ledger, current through 2026-08-17, aggregates Coding, Agentic, Reasoning, and Multimodal sources, but its total score and ranking are custom aggregates.Reddit LocalLLaMA's Experience with Kimi K2.5 Coding and DeploymentA Reddit LocalLLaMA thread focuses on Kimi K2.5 tool definitions, Agent Swarm, and deployment differences. Its alleged leaked prompt is incomplete and should be checked against official chat templates and provider behavior.Kimi K2.5 Official Release: Multimodality, Agent Swarm, and Coding BenchmarksKimi's official release positions K2.5 as a vision, coding, and Agent Swarm model and publishes Thinking, tool, context, and some benchmark conditions.Kimi K2.5 Vision Coding and Agent Swarm Task PromptKimi's official material connects K2.5 visual input, coding tasks, and Agent Swarm decomposition with completion checks and evidence capture.Kimi K2.5 Thinking/Instant and Vision Tool ConfigurationKimi's official repository separates Thinking and Instant parameters and vision-tool configuration, with chat-template and reasoning_content checks at deployment.