Gemma 4 × SillyTavern

Gemma 4 in SillyTavern

Trying Gemma 4 for character chat? First separate the official model names from community quant files. Then check the backend, chat template, tokenizer and thinking markers in that order. This guide gives you a short route to a working reply, plus a browser option when local setup is the part slowing you down.

Official Gemma 4 page
Verify names at Google DeepMind
Tabbit desktop new-tab window with vertical tabs, a central prompt and a chat panel on the right.

Start with the model

Gemma 4 is a family, not one SillyTavern preset

Google lists E2B, E4B, 12B, 26B A4B and 31B. The 26B A4B model is mixture-of-experts; 31B is dense. A Q4 or Q6 file is a quantized release, not a new official model.

01

26B A4B or 31B?

Use the model card and your hardware to decide. 26B A4B can be lighter to run because only part of the model is active. 31B usually asks more of memory. Do not compare the numbers as if they describe the same architecture.

02

Local, provider or API

Ollama, LM Studio, KoboldCpp and a llama.cpp server take different routes through templates and context. An OpenRouter or other API route may be simpler, but the exact model ID and alias belong to that provider.

03

Roleplay is a setup test

Character consistency depends on the card, system prompt, sampler and context budget as much as the checkpoint. Community reports mention less repetition with DRY and presets, but those are starting points, not official defaults.

SillyTavern path

A clean connection order

Change one layer at a time. That makes a bad reply diagnosable instead of turning every setting into a guess.

  1. 01

    1. Confirm the endpoint

    Pick the right connector for your backend and confirm that it lists the exact Gemma 4 checkpoint. If you use a provider, copy its model slug rather than typing a remembered alias.

  2. 02

    2. Choose Chat Completions

    For an API or local server that supports it, start with Chat Completions. Text Completion can work, but it exposes more template details and makes a stale format easier to miss.

  3. 03

    3. Match the tokenizer and template

    Use the Gemma 4 or Gemini tokenizer/template supplied by your backend. Older Gemma templates can put headers and reasoning markers in the wrong place.

  4. 04

    4. Add your card and preset

    Load the character card, then add a restrained system prompt. Try one community preset at a time. Keep the original files with their author and record what you changed.

The names FF, Moonlight and other presets come from community discussions. They are not Google defaults and this page does not republish their files.

When the reply looks wrong

Fix the layer that failed

CASE 01

Thinking text will not parse

Check the model template, tokenizer and thinking setting first. Some Reddit replies suggest a <|think|> marker in the system prompt, but that depends on the backend and can create literal tokens when the format is wrong.

CASE 02

It thinks only sometimes

A provider may expose reasoning differently from a local build. Do not force a marker blindly. Compare one short request with thinking on and off, then read the provider or model-card instructions.

CASE 03

The character repeats or drifts

Lower the context noise before changing temperature. Trim duplicated card text, check the context limit, then try a community sampler or DRY setting. Reddit users often mention Temperature 1.0, Top-K 64 and Top-P 0.95, but treat these as experiments.

A browser route

Keep the character page open and ask beside it

Tabbit does not replace SillyTavern cards or lorebooks. It solves a different part of the problem: reading a live wiki, document or reference page while chatting with a model already available in the browser.

  1. 1

    Install Tabbit

    Download the Chromium-based browser for macOS or Windows. No local model file or SillyTavern server is required for this route.

  2. 2

    Choose what is actually listed

    Open the model picker after install and select a current model shown there, such as GPT-5.4, GPT-5.2-Chat or Gemini-3-Flash. Availability can change, so the picker is the source of truth.

  3. 3

    Reference the live page

    Keep the character wiki or setting document in a tab. Use @ to reference the page, a screenshot or a file, then ask for a summary, scene notes or a second take.

Tabbit new-tab model picker showing GPT-5.4, GPT-5.2-Chat, Gemini-3.1-Pro, Gemini-3-Flash and Claude-Sonnet-4.6. Gemma 4 is not visible in this screenshot.

Pick the right tool

SillyTavern or Tabbit?

Use the one that matches the bottleneck. A browser shortcut is not a character-card manager, and a local frontend is not automatically a better model.

SillyTavernTabbit
Character cards and lorebooksBuilt inKeep the source page open
Gemma 4 local checkpointBring your own backendUse a model shown in the picker
Template and sampler controlDetailedManaged by the selected model
Page contextPaste or use extensions@ the live tab or file
First useful replyInstall + endpoint + presetInstall + choose a listed model
Tabbit new-tab model picker showing GPT-5.4, GPT-5.2-Chat, Gemini-3.1-Pro, Gemini-3-Flash and Claude-Sonnet-4.6. Gemma 4 is not visible in this screenshot.

Reference the live page

Keep the character wiki or setting document in a tab. Use @ to reference the page, a screenshot or a file, then ask for a summary, scene notes or a second take.

Tabbit new-tab model picker showing GPT-5.4, GPT-5.2-Chat, Gemini-3.1-Pro, Gemini-3-Flash and Claude-Sonnet-4.6. Gemma 4 is not visible in this screenshot.

Page context

@ the live tab or file

FAQ

Gemma 4 and SillyTavern, answered

Is Gemma 4 26B the same as 31B?+

No. Google describes 26B A4B as a mixture-of-experts model and 31B as dense. Choose from the model card and hardware you actually have.

Which Gemma 4 quantization should I download?+

There is no universal best file. Match Q4, Q5 or Q6 to available memory, context needs and the publisher’s model card. A quant file is not an official new Gemma name.

Why do thinking tokens show as plain text?+

The tokenizer, template, backend or provider may disagree about the format. Verify those pieces before adding markers such as <|think|>; a marker in the wrong template can make the output worse.

Can Tabbit run Gemma 4?+

Do not assume it. Open Tabbit’s current model picker after installation and use a model that is actually listed there. The screenshot on this page does not show Gemma 4.

Does Tabbit replace SillyTavern?+

No. SillyTavern remains the better fit for cards, lorebooks and detailed preset control. Tabbit is for chatting beside live web pages and files.

Stop debugging the wrong layer

If you need Gemma 4 locally, use the model card and check the template chain. If you need an answer beside a character wiki now, install Tabbit and choose a model from its current picker.

Available for macOS and Windows. The model list can change.

© 2026 Tabbit Browser. The AI-native browser that understands your context.