Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Kimi K3 · Media / benchmark · Personal experience

Single SVG pelican case: useful signal, not an agent benchmark

Simon Willison used OpenRouter and llm-openrouter to generate an SVG and describe the rendered image; one personal task.

Media / benchmarkPersonal experienceEdited 2026-09-20

Test conditions

Condition
Model: moonshotai/kimi-k3.
Condition
Client: OpenRouter + llm-openrouter.
Condition
Task: SVG pelican riding a bicycle; 95 input tokens and 16,658 output tokens.
Condition
Timing: article dated 2026-07-16; single observation.

Key data and applicable tasks

Original article

Simon Willison’s Weblog Subscribe Kimi K3, and what we can still learn from the pelican benchmark

Chinese AI lab Moonshot AI announced Kimi K3 this morning, describing it as their “most capable model to date, with 2.8 trillion parameters”. It’s currently available via their website and API, but an open weight release is promised “by July 27, 2026”.

Moonshot are calling this the first “open 3T-class model” (I guess they’re rounding 2.8 trillion up to 3 trillion), taking the crown from DeepSeek’s 1.6T v4 Pro. Their self-reported benchmarks have K3 mostly beating Claude Opus 4.8 max and GPT-5.5 high, while losing out to Claude Fable 5 and GPT-5.6 Sol.

A few highlights from the Artificial Analysis report on the model:

“On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5.” “Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers” “Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6.”

The model is also now the leading model on Arena.ai’s Frontend Code arena, surpassing even Claude Fable 5.

The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic’s Claude Sonnet series and making it the most expensive model released by a Chinese AI lab to date. This is a significant increase on their earlier models such as Kimi K2.6 at $0.95/$4. 2.8 trillion parameters is also more than twice the size of that 1T model.

But how does it pelican? #

I used OpenRouter (to avoid signing up for a Moonshot API key) with the llm-openrouter plugin to generate an SVG of a pelican riding a bicycle:

llm -m openrouter/moonshotai/kimi-k3 'Generate an SVG of a pelican riding a bicycle'

Here’s the transcript. It looks like this:

That pelican took 95 input tokens and 16,658 output tokens (13,241 were reasoning tokens), for a total cost of 25 cents!

Since K3 accepts image input I ran it against that rendered SVG above (with my alt text prompt) and got back (for 0.6 cents):

Cartoon illustration of a white pelican wearing a red scarf, riding a red bicycle along a gray road with white dashed lines; the pelican has a large orange beak and webbed orange feet pedaling, with white motion lines behind it; the background shows a light blue sky with white clouds, a yellow sun, two small black birds in flight, and green grass with tiny white flowers in the foreground

What can we learn from the pelican? #

My Generate an SVG of a pelican riding a bicycle test is 21 months old now. It was never a particularly great benchmark. It started out as a joke on how absurdly difficult it is to compare these models, but then for the first year it turned out to have a surprising correlation to how good the models actually were.

That connection has been mostly severed now. The GPT-5.6 and Claude Fable 5 pelicans are outclassed by GLM-5.2, and much as I love GLM I don’t think that’s a Fable-class model.

(I’m still not convinced that labs are training for the benchmark—if they were, I’d expect much better results. There’s a chance that Gemini has optimized for any combination of an animal on a vehicle though!)

The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.

So don’t go using pelicans to compare models!

All of that said, I still get a decent amount of value out of running the benchmark myself.

Firstly, it’s a forcing function for actually trying the model. If I show you a pelican, that means I’ve managed to run a prompt through it. If the model has an official API I’ll use that, if it’s open weight (and small enough to fit a 128GB M5 MacBook Pro) I’ll try running it on my own machine, usually via llama.cpp or LM Studio or Ollama. I’ll frequently use OpenRouter since that usually provides a proxy to an official API without me needing a new API key.

Most of my pelicans are generated using my LLM CLI tool, which helps encourage me to ensure the latest models are supported by that (via one of its plugins).

More importantly though, even the act of a single prompt to “Generate an SVG of a pelican riding a bicycle” can reveal interesting model characteristics.

Consider the result for Kimi K3 today. Running those simple prompts helped emphasize several points about the model.

It only has one reasoning effort right now, “max”—and it shows. The model consumed 13,241 reasoning tokens to output 3,417 tokens of response. This is expensive—the pelican cost 25 cents! How does the prompt “Generate an SVG of a pelican riding a bicycle” add up to 95 input tokens? OpenAI’s tokenizer counts 10, Anthropic’s counts 10 for Opus 4.6, 30 for Opus 4.7 and 25 for Sonnet 5/Fable 5. Prom pting “hi” to Kimi K3 counted 86 tokens, suggesting there may be an 85 token hidden system prompt. It refused to leak it though. Vision works well: the alt text it generated is very good.

K3 currently only has one thinking effort level, but I’ve been deriving quite a bit of value recently from running the same pelican prompt through different effort levels to get a quick idea for what impact those have. Here’s my matrix for the GPT-5.6 model family, for example.

Really though the main things I gain from the pelican test are:

It’s a “hello world” exercise for prompting a model A rough cost and reasoning estimate for a simple task Confirmation that the model can output valid SVG and has a basic idea of geometry and spatial awareness. This is a much bigger deal for the smaller models that run on my laptop. It’s still interesting to compare pelicans between releases in the same model family. K3’s pelican is a notable improvement from Kimi 2.5. It’s something I can share that demonstrates I’ve tried it. Plus a comment with a pelican in it is kind of a tradition on Hacker News at this point, any time I’m late I get comments asking where it is! Posted 16th July 2026 at 8:19 pm · Follow me on Mastodon, Bluesky, Twitter or subscribe to my newsletter More recent articles Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026 One-shotting a Raccoon Heist game using Claude Fable 5 - 5th August 2026

This is Kimi K3, and what we can still learn from the pelican benchmark by Simon Willison, posted on 16th July 2026.

ai 2,189 generative-ai 1,939 llms 1,906 llm-pricing 89 pelican-riding-a-bicycle 135 llm-release 225 ai-in-china 105 artificial-analysis 8 moonshot 9 kimi 13

Next: A Fireside Chat with Cat and Thariq from the Claude Code team

Previous: The new GPT-5.6 family: Luna, Terra, Sol

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

Pay me to send you less!

Sponsor & subscribe Disclosures Colophon © 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026

undefined

What this supports

  • Simon Willison used OpenRouter and llm-openrouter to generate an SVG and describe the rendered image; one personal task.

What this does not support

  • Does not support inferring visual quality, cost, or stability from the single 2026-07-16 task with 95 input and 16,658 output tokens.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Google / Simon Willison · Author not disclosed · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Kimi K3

Compare Kimi K3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

Kimi K3 Pricing: API Costs, Subscriptions, and Budget Math

Kimi K3 pricing explained with official API rates, cache-write rules, current membership tiers, worked costs, and a practical choice framework.

Related reviews

Visual coding case study is useful, not a general success ratePuter Developer’s visual-coding case with target/current renders and code examples.Kimi K3 code security evaluation: strong benchmarks do not guarantee precisionSemgrep’s IDOR code-security benchmark, separating precision, recall, and F1.Real-world coding: close on simple tasks, weaker on trap tasksCoding-agent observation using identical GitHub issues, plan-build-validate stages, and a 70-point rubric.Coding-agent evidence is serious, but not “best overall”NxCode’s synthesis of public coding-agent benchmarks, configurations, and comparability limits.Configure Kimi K3 developer calls and multimodal inputsA sourced guide for configure kimi k3 developer calls and multimodal inputs, with explicit inputs, environment, and boundaries; see the detail page for the execution path.Generate a checkable project scaffold from one task briefA sourced guide for generate a checkable project scaffold from one task brief, with explicit inputs, environment, and boundaries; see the detail page for the execution path.Use Kimi K3 with OpenCode and Firecrawl for sourced web researchConnect Kimi K3, OpenCode, and Firecrawl MCP into a cited web-research workflow with explicit domains, permissions, and stop rules.Write Kimi API requests as testable tasksTurn the official prompting guidance into a checklist for role, context, constraints, format, and acceptance. The source does not provide one complete reusable prompt.