Use the official checkpoint
For the instruction variant, start from a provider or repository that identifies it as Gemma 4 31B IT. GGUF names from Unsloth, bartowski or LM Studio Community are quantized distributions, not new Google model families.
Gemma 4 31B × SillyTavern
The 31B query hides the first trap: Google calls this a dense Gemma 4 model, while 26B A4B is a different MoE model. This guide maps the name, memory, quantization, provider and chat template before you tune a character card.
Verify the official naming and template ↗
Name it correctly
Google lists Gemma 4 31B as a dense model with about 30.7B parameters and a 256K context window. Gemma 4 26B A4B has about 25.2B total parameters but about 3.8B active parameters because it is MoE.
For the instruction variant, start from a provider or repository that identifies it as Gemma 4 31B IT. GGUF names from Unsloth, bartowski or LM Studio Community are quantized distributions, not new Google model families.
A Q4 file is smaller than full precision, but the model still has 31B active parameters. Community reports put Q4_K_M around 18 GB for weights, before context, runtime overhead and any multimodal files. Treat that as a planning figure, not a guarantee.
SillyTavern users report strong creative writing from both base and instruct files, but they need different prompts and templates. Decide whether you want instruction following or a base checkpoint before comparing presets.
SillyTavern path
Change one layer at a time. That makes a bad reply diagnosable instead of turning every setting into a guess.
Pick the right connector for your backend and confirm that it lists the exact Gemma 4 checkpoint. If you use a provider, copy its model slug rather than typing a remembered alias.
For an API or local server that supports it, start with Chat Completions. Text Completion can work, but it exposes more template details and makes a stale format easier to miss.
Use the Gemma 4 or Gemini tokenizer/template supplied by your backend. Older Gemma templates can put headers and reasoning markers in the wrong place.
Load the character card, then add a restrained system prompt. Try one community preset at a time. Keep the original files with their author and record what you changed.
The names FF, Moonlight and other presets come from community discussions. They are not Google defaults and this page does not republish their files.
When the reply looks wrong
Check the model template, tokenizer and thinking setting first. Some Reddit replies suggest a <|think|> marker in the system prompt, but that depends on the backend and can create literal tokens when the format is wrong.
A provider may expose reasoning differently from a local build. Do not force a marker blindly. Compare one short request with thinking on and off, then read the provider or model-card instructions.
Lower the context noise before changing temperature. Trim duplicated card text, check the context limit, then try a community sampler or DRY setting. Reddit users often mention Temperature 1.0, Top-K 64 and Top-P 0.95, but treat these as experiments.
A browser route
Tabbit does not replace SillyTavern cards or lorebooks. It solves a different part of the problem: reading a live wiki, document or reference page while chatting with a model already available in the browser.
Download the Chromium-based browser for macOS or Windows. No local model file or SillyTavern server is required for this route.
Open the model picker after install and select a current model shown there, such as GPT-5.4, GPT-5.2-Chat or Gemini-3-Flash. Availability can change, so the picker is the source of truth.
Keep the character wiki or setting document in a tab. Use @ to reference the page, a screenshot or a file, then ask for a summary, scene notes or a second take.

Pick the right tool
Use the one that matches the bottleneck. A browser shortcut is not a character-card manager, and a local frontend is not automatically a better model.
| SillyTavern | Tabbit | |
|---|---|---|
| Character cards and lorebooks | Built in | Keep the source page open |
| Gemma 4 local checkpoint | Bring your own backend | Use a model shown in the picker |
| Template and sampler control | Detailed | Managed by the selected model |
| Page context | Paste or use extensions | @ the live tab or file |
| First useful reply | Install + endpoint + preset | Install + choose a listed model |

Keep the character wiki or setting document in a tab. Use @ to reference the page, a screenshot or a file, then ask for a summary, scene notes or a second take.

@ the live tab or file
FAQ
No. Google describes 26B A4B as a mixture-of-experts model and 31B as dense. Choose from the model card and hardware you actually have.
There is no universal best file. Match Q4, Q5 or Q6 to available memory, context needs and the publisher’s model card. A quant file is not an official new Gemma name.
The tokenizer, template, backend or provider may disagree about the format. Verify those pieces before adding markers such as <|think|>; a marker in the wrong template can make the output worse.
Do not assume it. Open Tabbit’s current model picker after installation and use a model that is actually listed there. The screenshot on this page does not show Gemma 4.
No. SillyTavern remains the better fit for cards, lorebooks and detailed preset control. Tabbit is for chatting beside live web pages and files.
If you need Gemma 4 locally, use the model card and check the template chain. If you need an answer beside a character wiki now, install Tabbit and choose a model from its current picker.
Available for macOS and Windows. The model list can change.