This discussion offers no single answer: CptPhantasmic found V4.1 Flash promising for writing style, instruction following, and detail handling in a short Janitor test, while MikabellStarfall encountered single replies exceeding 2,000 words, loss of character control, and forgetting. Both emphasize the need to adjust the prompt or character card.
Tasks it is suitable for assessing: Writing style, prompt adaptation, instruction following, character control, and plot continuity risks in Janitor roleplay.
Tasks it is unsuitable for extrapolating to: General model quality, a strict verdict on whether V4 Pro is better or worse, or consistent performance across all clients.
Applicable model version: DeepSeek V4.1 Flash.
Test environment or client: Both CptPhantasmic and MikabellStarfall mention Janitor; CptPhantasmic also says they plan to test a modified FF4 MAX+ prompt in SillyTavern (ST). The provider, full preset, and API configuration are not stated.
Reasoning tier and parameters: Not stated.
The main post only asks whether V4.1 Flash is better than V4 Pro; it defines no task set. CptPhantasmic says they ran a quick test on a small number of bots in Janitor, while MikabellStarfall describes an experience using a starter prompt. Comments also suggest adjusting the preset, prompt, and character card for each model. There was no uniform prompt, sample size, number of repetitions, baseline, or independent review, so this can only be treated as community opinion.
Positive experience: CptPhantasmic says that after tweaking the prompt, the model performed “quite well.” Its writing style differs from Pro-0813, it can adapt to different types, and it follows instructions and details relatively well. He says it did not skip or skim prompts and replies as V4 Flash did, and he did not encounter common DeepSeek stock phrasing. This conclusion comes from a small, quick Janitor test and does not cover deep plotlines or positivity bias.
Negative experience: MikabellStarfall says Janitor turns the opening prompt into a short story of 2,000+ words, follows the opening only loosely, and adds setting, character details, and plot. Even when the prompt says not to take over the user's turn, it takes over the entire scene, making back-and-forth co-writing difficult. She also describes the model as very forgetful: the prompt specified the Grand Pas de Deux from The Nutcracker, but the model wrote the Sugar Plum Fairy instead, and in the same message regressed to Act I.
Other signals: Kahvana says the preset and character card should be adjusted for the new model; PhysicalKnowledge also says the prompt needs slight changes. ForwardMastodon considers it worse than V4 Flash; SleepBaobei reports frequent logical errors and hallucinations in simple plots. These opinions conflict and cannot be combined into a single score.
| Observer | Client and scope | Observed phenomenon |
|---|---|---|
| CptPhantasmic | Janitor; quick test on a small number of bots | Different writing style, adaptation to different types, and relatively good instruction and detail following; prompt still needs fine-tuning |
| MikabellStarfall | Janitor; a single starter prompt | 2,000+ words, extra details, taking over the scene, and continuity and memory problems |
| Kahvana, PhysicalKnowledge | Client not fully specified | Suggested adjusting the preset, character card, or prompt |
| Thread snapshot | 5 upvotes, 16 comments | Discussion size, not an evaluation score |
This post supports a narrow conclusion: the roleplay experience with V4.1 Flash may depend heavily on Janitor's preset, character card, and prompt design. It may improve writing style and instruction following, or it may disrupt co-writing through overgeneration. MikabellStarfall's The Nutcracker mix-up is a single continuity case and is insufficient to prove a general memory defect; CptPhantasmic's positive conclusion also covers only a short test. The original post does not disclose the full preset or prompt, so it cannot support a claim that a directly copyable configuration exists.
Fix the same Janitor character card, preset, opening prompt, temperature, and other parameters, then run V4.1 Flash alongside a comparison model with an explicit version. Repeat the test while recording reply word count, extra facts, whether the model takes over the user's character, instruction following, and cross-turn plot consistency. If using the modified FF4 MAX+ prompt in ST, obtain the complete text and record the configuration; the post itself does not provide them, so it cannot be directly reproduced from this material.
DeepSeek V4.1 Flash