A one-week roleplay test found that LongCat 2.0 was "a jackpot" for creative writing: faithful instruction following, coherent stories, no hard refusals, and dry, non-sensational narration. It also had two clear flaws — excessive fidelity to supplied information leading to repetitive inertia, and the need for extremely specific requests to reach jailbreak content (the author had to use an exceptionally strong jailbreak prompt).
Faithful and accurate instruction following: it can gracefully handle very complex worldbuilding setups.
Coherent stories: in every scene, the author never encountered text that felt broken or garbled.
Zero hard censorship: it never produces "I'm sorry, but I cannot..." for content of any kind; the author tested scenes that most people would not be able to keep reading, and the model offered "absolute freedom."
No exaggerated or sentimental narration: the prose stays dry and neutral and does not force a narrator's perspective into the story (the author's comparison: the GLM series is full of literary narrator's voice, which is why they dislike GLM).
Surprisingly knowledgeable about niche preferences: with correct instructions, it can reproduce extremely fine details from adult media.
(Key) Too faithful to supplied information: unless long stories are aggressively protected against repetition, nearly identical phrases and scenes may recur; it has a strong "copy-and-paste" inertia for content written into the worldbuilding and instructions — even low-priority instructions can seep in unexpectedly, so worldbuilding and instructions must be designed carefully.
(Key) Strange censorship mechanism: extreme content is not refused, but you must be very specific about what you want — "It's like a library: usually you have to wrestle with the librarian's access controls; with Longcat, the books of extreme content are simply deep in the hallway, with nothing stopping you, but reaching them requires an effortful stretch." The author therefore had to use an exceptionally strong jailbreak prompt.
The repetition issue varies across providers (the author canceled their NanoGPT subscription and switched to OpenRouter as the primary provider).
The author used the non-reasoning / non-thinking version.
A comment added that Owl Alpha and Longcat 2.0 are the only models able to distinguish Algerian/Moroccan darija dialects and understand certain nearly map-disappeared coastal towns, regional cultures, and regional conflicts (a North African dialect roleplay scenario on r).
This is a single anonymous user in a single setting (SillyTavern roleplay, non-thinking mode), so the conclusion does not represent all tasks. "Zero censorship" occurred under that mode and prompt setup; comments also discuss censorship differences between the official API (which requires phone-number registration) and OpenRouter, with no consensus.
Where it aligns with the official positioning: strong instruction following and good Agent/long-range task performance (official SWE-bench, etc.); where it differs from official promotion: the official source does not mention the friction of needing extremely specific requests for censored content.
Task fit: suitable for creative writing, RP, and long-form narratives (add your own anti-repetition instructions); unsuitable for compliance scenarios that require explicit censorship boundaries.
Related: another r/SillyTavernAI thread (1u4ts39) reports that during multi-character RP it "likes to insert characters who are not present," which can serve as an additional negative observation for multi-character scenarios.
LongCat 2.0