SETUP BEFORE PRESETS

Kimi K3 setup in SillyTavern

A working connection starts with four details: the live provider, the model alias, the Chat Completions path, and how thinking history is handled. Check those before tuning a character preset.

Jump to the checklist

Official alias: kimi-k3 · Example base URL: https://api.moonshot.ai/v1 · K3 always has thinking enabled.

Tabbit desktop new-tab screen with vertical tabs, a centered input, and a Chat panel.

WHAT THE SIGNALS MEAN

Separate model facts from setup noise

Kimi K3 can be a good writing model and still expose a bad endpoint or an incomplete reasoning message. These three checks keep the diagnosis honest.

01 / MODEL

Thinking is always on

Moonshot documents K3 as an always-thinking model. The API accepts reasoning_effort low, high, or max. A long reply is not automatically a broken preset.

02 / PROVIDER

The route controls the limits

The official example uses https://api.moonshot.ai/v1 and model kimi-k3. Proxies can expose different aliases, quotas, prices, or parameter support.

03 / REPORTS

429 and slow replies are reports

Public X and Reddit discussions mention slow responses, 429s, overthinking, and reasoning blocks. They describe experiences, not a service-level promise.

THE FIVE-CHECK SETUP

Connect Kimi K3 one layer at a time

Use a small test chat first. Change one setting, send one turn, and keep a note of what changed.

K3 API facts come from the official Kimi guide. Provider dashboards and SillyTavern extension behavior can change independently.

  1. 01

    Verify the provider and model list

    Choose Moonshot or a provider that currently lists Kimi K3. Confirm billing, rate limits, the live model slug, and whether the route is OpenAI compatible.

    You can point to the provider entry that says Kimi K3 or kimi-k3.

  2. 02

    Use the Chat Completions connector

    The official Quickstart sends requests to /chat/completions. In SillyTavern, use the Chat Completions path rather than a legacy text-completion template.

    The connection test reaches /v1/chat/completions without a template error.

  3. 03

    Start with a short system prompt

    Use one character, one scene, and a small context window. State voice and boundaries once. Contradictory patches make it harder to tell whether the API works.

    The first reply has the expected role and format before any preset is added.

  4. 04

    Map reasoning and preserve the turn

    K3 streaming can return reasoning_content separately from content. Make sure the frontend handles reasoning blocks and sends the complete assistant message on the next turn.

    The next request does not silently drop the assistant reasoning or final message.

  5. 05

    Add prefill or a preset last

    Prefill and preserved-thinking extensions can help with a specific RP behavior, but they are community tooling, not a Kimi requirement. Change one layer and keep a rollback copy.

    You can name the one change that improved or worsened the reply.

SYMPTOM FIRST

Fix the layer that is actually failing

Use the narrowest check that matches the symptom. Do not rotate through random presets while an auth or endpoint error remains.

Fix the layer that is actually failing
What you seeLikely layerNext check
401 or "invalid key"Key or accountCreate a fresh key in the provider console, confirm the account is unlocked, and remove stray spaces.
404 or model not foundBase URL or aliasCompare the live provider model list with kimi-k3. Keep /v1 in the base URL and let the connector add /chat/completions.
429, timeout, or “too busy”Quota or rate limitCheck concurrency, RPM/TPM, account tier, and provider status. Retry only after confirming the request is not looping.
Empty reply or broken reasoning blockResponse mappingInspect reasoning_content versus content, enable the connector reasoning option, and preserve the complete assistant message.
Long thoughts, drift, or repetitionPrompt and historyLower reasoning_effort when the provider supports it, shorten the instruction stack, and test a clean chat before changing prefill.

A SHORTER PATH

Use Kimi K3 beside the context you already have

If the immediate job is reading a character sheet, wiki, or writing brief while chatting, Tabbit removes the first endpoint setup loop. Check the live picker for access and availability.

01

Keep the source page open

Open the character sheet, lore page, or reference document in a tab. Tabbit can use visible pages, screenshots, and files as context.

The source stays visible while you ask the question.

Tabbit Agent view operating in a Google Sheet with a task prompt and execution steps on the right.
02

Choose Kimi-K3 in the picker

Select Kimi-K3 from the current Chat or new-tab model picker. The model roster can change, so treat the live UI as the source of truth.

The picker shows the model you intend to use.

Tabbit sidebar beside an article with summary, Switch Model, and a question field.
03

Ask, compare, and continue reading

Use the sidebar for a focused question or compare model replies in one view. Tabbit is a browser chat route, not a replacement for ST cards, lorebooks, or group chat.

Your prompt keeps the page or file context in the same workflow.

Tabbit five-column model chat with Kimi-K3 visible in the middle column.

SETUP FAQ

Answers before you change another knob

What is the official Kimi K3 API alias?+

The official Quickstart uses model kimi-k3 and the example base URL https://api.moonshot.ai/v1. Verify the live provider list because aliases and availability can change.

Which SillyTavern connection type should I use?+

Use an OpenAI-compatible Chat Completions connection for the official Moonshot route. A text-completion template is a different request shape.

Why does Kimi K3 think for so long?+

K3 always has thinking enabled in the Moonshot API. Where supported, try a lower reasoning_effort, keep instructions short, and confirm the full assistant message is preserved.

Do I need a Kimi K3 preset or prefill extension?+

No. Start with the clean connection first. Add community preset or prefill guidance only when you can describe the behavior you want to change.

What should I do about 429 errors?+

Treat 429 as a quota or rate-limit signal. Check account tier, concurrency, RPM/TPM, provider status, and whether SillyTavern is retrying repeatedly.

Does Tabbit replace SillyTavern?+

No. Tabbit is useful for using Kimi K3 beside a live page or file. SillyTavern remains the tool for its character cards, lorebooks, extensions, and group chat.

Test the model before tuning the whole stack

Keep SillyTavern for its character workflow. For a lower-friction first chat, open Tabbit, choose Kimi-K3, and bring the page context with you.

Available for macOS and Windows. Model access and quotas may vary by edition and plan.

© 2026 Tabbit Browser. The AI-native browser that understands your context.