Mistral Small 3.2 24B
Official instruct checkpoint
Best baseline for precise instructions, multilingual chat and a clean Apache-2.0 starting point.
General-purpose behavior may feel less literary or proactive than an RP tune.
Open model card ↗MISTRAL 24B / RP FIELD GUIDE
The right pick changes with your character card, hardware and taste. Start with the exact repository ID, then compare voice, character consistency, initiative, repetition, context and license. This shortlist separates official Mistral facts from community reports.

CONDITIONAL VERDICT
These are practical starting points, not a universal leaderboard. Model cards and community comments describe different tests, so treat each fit as a hypothesis to verify with your own card.
Official instruct checkpoint
Best baseline for precise instructions, multilingual chat and a clean Apache-2.0 starting point.
General-purpose behavior may feel less literary or proactive than an RP tune.
Open model card ↗Community RP / Magistral-style tune
Try it when prose, initiative and flexible system prompts matter more than conservative refusal behavior.
A community tune with its own template and sampler advice. Do not treat quoted praise as a benchmark.
Open model card ↗Community RP / Mistral Small 3.2 tune
A sensible test for character voice, group chat and scene development with a balanced writing style.
The model card reports occasional physics or causality slips. The Heretic variant is a separate artifact.
Open model card ↗Roleplay-first merged LoRA
Pick it for immersive narration, strong character voice and scene momentum.
The card lists long-generation repetition, user-character hijacking and weaker strict multi-character formatting.
Open model card ↗RP fine-tune with GGUF files
Useful when your local stack already favors llama.cpp or GGUF and you want clear file-size choices.
Q4, Q6 and Q8 files trade memory, speed and fidelity. The quant page is not a new official release.
Open model card ↗DECISION RUBRIC
Score the dimensions that change your actual sessions. A model that writes beautiful first turns but loses a minor character detail at 16k may be worse for your campaign than a plainer model that remembers.
| Criterion | Official 3.2 | Community RP | Best fit |
|---|---|---|---|
| Character consistency | Reliable instruction baseline | Often stronger voice, verify over 8 to 16k turns | Long-running cards |
| Prose and initiative | Controlled and direct | Magidonia, Cydonia and Umbra target creative momentum | Literary scenes |
| Repetition | 3.2 reports 2x fewer infinite generations than 3.1 | Depends on tune, quant and sampler | Long outputs |
| Context behavior | 3.2 card lists a 128k-capable family | Practical KV cache still depends on backend and quant | Lore-heavy chats |
| Hardware | About 55GB GPU memory in BF16/FP16 | GGUF examples range from 13.6GB Q4_K_S to 25.2GB Q8_0 | Local inference |
| License and freshness | Apache 2.0; official release ID is clear | Check each card, base, date and variant | Sharing or shipping |
Choice score = 0.30 × card fit + 0.20 × voice + 0.20 × consistency + 0.15 × context + 0.10 × hardware fit + 0.05 × license confidence. Change the weights when your use case changes.
A REPEATABLE TEST
Do not compare a fresh base model against a tuned model with different prompts. Keep the same character, scene seed, context length and output budget. Save the raw prompt and the quant file beside every result.
Give the same card, setting and opening action to each candidate. Check diction, point of view, initiative and whether the model respects the user character boundary.
Add a small fact, a new emotional beat and a second character. Mark forgotten details, repeated phrases, sudden tone changes and uninvited actions.
Rate consistency, prose, agency, repetition and latency from 1 to 5. Record temperature, top_p, context, quant and backend so another run means something.
A winning candidate is the one that fits your card and machine after fixes. Re-test a new quant or template before discarding a model.

LOCAL SETUP CHECK
A weak-looking reply often starts one layer earlier than the model weights. Confirm the exact ID, format and chat template before you tune generation.
3.2, 3.1, base, instruct, Magidonia, Cydonia and Umbra are separate artifacts. A label such as "Mistral 24B" is not enough.
The 3.2 card cites about 55GB GPU memory for BF16/FP16. GGUF sizes vary by quant. Leave room for KV cache, context and the frontend.
Mistral cards point to Mistral tokenizer and chat formatting. If `<s>`, role markers or stop strings leak into the reply, remove duplicate template handling.
The official card suggests low temperature for general use. RP cards may expect a different range. Change one value and keep the prompt fixed.
TABBIT AS THE RESEARCH DESK
Tabbit does not run Mistral 24B and does not replace your local RP frontend. The public model directory checked for this guide did not show Mistral. Use Tabbit to keep official cards, quant pages, Reddit notes and your test results in one browser context.
Put exact IDs, license notes and update dates in adjacent tabs. Separate evidence from a quoted user preference.
Bring a model card, screenshot or local test note into a question and ask for a checklist that preserves IDs and labels unknowns.
Tabbit can compare multiple visible model responses and summarize sources. Its picker changes, so check the live list after installation.


Model desk
IDs stay visible. Sources stay close. Your browser becomes the lab notebook.
Use the source cards above as the record of what was checked, not as a promise that every endpoint or model remains available.
FAQ
There is no source-backed universal winner. Start with Mistral Small 3.2 for a clean baseline, then test Magidonia, Cydonia, Umbra or Roleplay V3 against your own card and hardware.
Search results identify Skyfall as a 36B upscale and a 31B community model. It may descend from Mistral Small 24B, but it is not a 24B checkpoint. Compare by exact repository ID.
Community cards report strong consistency for Magidonia and Cydonia, while Umbra focuses on character voice and scene momentum. These are qualitative reports, so run the two-turn test.
Use the highest-quality quant your memory and context budget can sustain. The Roleplay V3 GGUF page lists Q4_K_S at 13.6GB, Q4_K_M at 14.4GB, Q6_K at 19.4GB and Q8_0 at 25.2GB.
Do not assume it. Mistral was not visible in Tabbit's public model directory when this guide was checked. Tabbit is recommended here for source collection and comparison.
Keep the exact checkpoint, quant, prompt and result together. That makes a new model release easier to compare and a bad session easier to diagnose.
Available for macOS and Windows. Tabbit model availability can change.