Qwen3.8-27B · SillyTavern field guide

Qwen3.8-27B in SillyTavern

The name is real, but it is easy to mix up. Qwen publishes an official Qwen3.8-27B dense vision-language checkpoint. Qwen3.8-Max is a different family member: 2.4T parameters in total and 95B active. Start with the checkpoint and provider you actually have, then tune the RP preset.

Verify the model identity

This guide separates upstream facts from community reports. It does not promise a specific provider, quantization or Tabbit account access.

Tabbit new tab showing the model picker and an @ reference field

Name check

27B is not the 2.4T model

Search results put several labels next to each other. They point to different artifacts, so a copied name can send a local setup down the wrong path.

01

Official dense checkpoint

The Qwen model card is Qwen/Qwen3.8-27B. It reports 27B parameters, a vision encoder, Apache-2.0 licensing and native 262,144-token context.

02

Qwen3.8-Max family

Qwen’s release post and related open-weight naming describe Qwen3.8-Max as 2.4T total with 95B active parameters. Its open-weight model is Qwen3.8-2.4T-A95B, not a 27B alias.

03

Community files

GGUF, FP8, GPTQ and roleplay fine-tunes can add a suffix or a creator name. Treat each repository as its own artifact and read its quantization and template notes.

04

Hosted IDs

SillyTavern needs the exact model ID exposed by the provider. A friendly display name is not proof that the route serves Qwen3.8-27B.

Model card

What the official 27B page actually says

Use these facts to choose a runtime. They describe the upstream checkpoint, not the memory cost of every quantized file.

27B parameters

The dense language model is listed at 27B parameters. The page also lists a 28B safetensors file size label, so do not turn a download label into a different model name.

Vision and video

The model card labels it image-text-to-text and describes native image and video understanding. A backend still needs the matching processor and multimodal support.

262K context

Native context is 262,144 tokens, extensible to 1M in the card. Long context does not mean a consumer GPU can load every configuration.

Serving choices

The card gives examples for Transformers, vLLM, SGLang, TokenSpeed and Docker. Choose a current version that explicitly supports this architecture.

Tabbit Deep Research page with Google results and execution steps for checking model documentation

SillyTavern setup

Work from the endpoint inward

A character card cannot fix a wrong route or a duplicated template. Use this order for a hosted API or a local server.

Provider labels, community presets and quantized repositories are not part of the upstream model card. Keep those claims attached to their source.

  1. 1. Name the artifact

    Copy the exact repository or provider model ID. Decide whether you have Qwen/Qwen3.8-27B, a quantized derivative or the separate Qwen3.8-2.4T-A95B family member.

  2. 2. Check the backend

    For local use, confirm that Transformers, vLLM, SGLang, llama.cpp or your chosen app supports the checkpoint and its vision path. A GGUF file may not expose every multimodal feature.

  3. 3. Connect SillyTavern

    For an OpenAI-compatible server, select Chat Completions, enter the documented base URL and copy the model ID returned by the server. Keep keys out of character cards.

  4. 4. Own the template once

    Let the hosted route serialize messages, or use the local tokenizer’s documented chat template. Applying both can produce broken roles, stray tags or an oddly terse reply.

  5. 5. Compare thinking

    Run the same card with thinking disabled or low, then medium and xhigh when the provider exposes those controls. Record latency and scene movement before changing the prose preset.

Fast diagnosis

Fix the boring failure first

Hold the character card and opening message steady. Change one input at a time, and save the returned model ID with your test notes.

SymptomCheck firstSmall fix
Model not found or silent fallbackProvider model list and response modelCopy the exposed ID. Remove a guessed qwen3.8-27b alias.
Roles, tags or formatting breakWho owns the chat templateUse one template path only, then start a fresh chat.
VRAM runs outWeight format, context and vision inputsTry a smaller quant, lower context and text-only baseline before blaming RP.
Long thinking, little story progressReasoning mode, output cap and preset lengthCompare low or medium with xhigh on the same card.
Prose changes between providersProvider filters, sampler defaults and system promptRecord the route as the variable. Do not call one provider result a model fact.

RP test card

Give the model a short contract

Use a compact test prompt while debugging. It makes template, thinking and provider differences easier to see.

Role: You are [character].
Scene: [place, relationship, current conflict].
Style: [POV, language, sentence length].
Direction: Advance one observable action. Do not speak for the user.
Continuity: Use confirmed facts only; ask when uncertain.
Format: Action first, dialogue second. 250 to 450 words. No recap.
Check: Keep the voice, boundaries and unresolved clue from the previous turn.

Send the same card three times. Compare latency, unwanted tags, format adherence and whether the scene changes. Only then adjust temperature, context or a community preset.

Browser route

Keep model docs and lore in the same workspace

SillyTavern remains the specialist cockpit for cards and lorebooks. Tabbit helps when the hard part is reading a model card, comparing provider notes and writing against a reference without constant copy and paste.

  1. 1

    Open the source

    Keep the Hugging Face card, provider docs or character wiki in a tab. The page stays available while you work through the setup.

  2. 2

    Check the live picker

    Open Tabbit’s model picker and choose Qwen3.8-27B only when your edition and account show that exact option. The picker is the availability source for Tabbit.

  3. 3

    Reference with @

    Use @ to bring a page, screenshot or file into a question. Ask another model to extract template requirements or compare a quantization note with the model card.

  4. 4

    Compare the answer

    Multi-model chat can read the same source from more than one angle. The interface proves the browser workflow, not identical model access for every user.

Tabbit page summary sidebar keeping a reference article and AI notes visible together

Browser route

Compare the answer

Tabbit page summary sidebar keeping a reference article and AI notes visible togetherTabbit Deep Research page with Google results and execution steps for checking model documentation

Choose the surface

SillyTavern or Tabbit?

They solve different parts of a roleplay workflow. Keep the specialist controls where they help, and use the browser when references are the bottleneck.

NeedSillyTavernTabbit
Cards and lorebooksDedicated RP controlsReference pages and files
Preset and prompt orderFine-grained prompt controlsSystem prompt and Skills
Read docs beside RPExtensions or copy and pasteOpen tabs become context
Compare model answersProvider setup or extensionsMulti-model browser chat
First actionVerify ID, backend and templateOpen the source, then check the live picker

FAQ

Qwen3.8-27B SillyTavern questions

Is Qwen3.8-27B an official model?+

Yes. Qwen/Qwen3.8-27B has an official Hugging Face model card. It is a 27B dense vision-language checkpoint. It is separate from Qwen3.8-Max and Qwen3.8-2.4T-A95B.

Is 27B the same as Qwen3.8-Max?+

No. Qwen3.8-Max is described as 2.4T total parameters with 95B active. Do not use a Max provider ID or a 2.4T quantization when you intend to test the 27B dense model.

Which quantization should I download?+

There is no single answer without your GPU, RAM, context target and backend. Start from a repository that states its format and loader support. Treat GGUF, FP8, GPTQ and fine-tune names as separate artifacts.

Which chat template belongs in SillyTavern?+

A hosted provider normally owns serialization. A local server should use the tokenizer or checkpoint documentation. Do not apply a second Qwen template on top of an already formatted endpoint.

Why is roleplay slow or too short?+

Check reasoning mode, output cap, context size, sampler defaults and provider routing separately. Community reports vary, so reproduce the same card and record the route before changing the prompt.

Can Tabbit run Qwen3.8-27B?+

Only choose it if your Tabbit model picker shows the exact model. Availability varies by edition, region, account and rollout. You can still use Tabbit to keep the model card or character reference beside a supported chat.

Leave the reference open

Install Tabbit for macOS or Windows, check the live model picker, and bring the model card or character wiki into the same browser conversation.

Model access and provider terms vary by edition and rollout.

© 2026 Tabbit Browser. The AI-native browser that understands your context.