Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

LongCat 2.0 · Community source · Personal experience

r/SillyTavernAI Field Test: One Week of LongCat 2.0 Roleplay (Writing/Jailbreaks/Repetition Tendency)

A one-week roleplay test found that LongCat 2.0 was "a jackpot" for creative writing: faithful instruction following, coherent stories, no hard refusals, and dry, non-sensational narration. It also had two clear flaws — excessive fidelity to supplied information leading to repetitive inertia, and the need for extremely specific requests to reach jailbreak content (the author had to use an exceptionally strong jailbreak prompt).

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model/version
LongCat-2.0; source date: 2026-08-18.
Harness/task
Pros; Faithful and accurate instruction following: it can gracefully handle very complex worldbuilding setups.
Sample/gaps
Limitations noted: Task fit: suitable for creative writing, RP, and long-form narratives (add your own anti-repetition instructions); unsuitable for compliance scenarios that require explicit censorship boundaries.; Related: another r/SillyTavernAI thread (1u4ts39) reports that during multi-character RP it "likes to insert characters who are not present," which can serve as an additional negative observation for multi-character scenarios.

Key data and applicable tasks

One-sentence takeaway

A one-week roleplay test found that LongCat 2.0 was "a jackpot" for creative writing: faithful instruction following, coherent stories, no hard refusals, and dry, non-sensational narration. It also had two clear flaws — excessive fidelity to supplied information leading to repetitive inertia, and the need for extremely specific requests to reach jailbreak content (the author had to use an exceptionally strong jailbreak prompt).

Original report (key points from the original)

Pros

  1. Faithful and accurate instruction following: it can gracefully handle very complex worldbuilding setups.

  2. Coherent stories: in every scene, the author never encountered text that felt broken or garbled.

  3. Zero hard censorship: it never produces "I'm sorry, but I cannot..." for content of any kind; the author tested scenes that most people would not be able to keep reading, and the model offered "absolute freedom."

  4. No exaggerated or sentimental narration: the prose stays dry and neutral and does not force a narrator's perspective into the story (the author's comparison: the GLM series is full of literary narrator's voice, which is why they dislike GLM).

  5. Surprisingly knowledgeable about niche preferences: with correct instructions, it can reproduce extremely fine details from adult media.

Cons

  1. (Key) Too faithful to supplied information: unless long stories are aggressively protected against repetition, nearly identical phrases and scenes may recur; it has a strong "copy-and-paste" inertia for content written into the worldbuilding and instructions — even low-priority instructions can seep in unexpectedly, so worldbuilding and instructions must be designed carefully.

  2. (Key) Strange censorship mechanism: extreme content is not refused, but you must be very specific about what you want — "It's like a library: usually you have to wrestle with the librarian's access controls; with Longcat, the books of extreme content are simply deep in the hallway, with nothing stopping you, but reaching them requires an effortful stretch." The author therefore had to use an exceptionally strong jailbreak prompt.

  3. The repetition issue varies across providers (the author canceled their NanoGPT subscription and switched to OpenRouter as the primary provider).

Important environment note (original post EDIT)

  • The author used the non-reasoning / non-thinking version.

  • A comment added that Owl Alpha and Longcat 2.0 are the only models able to distinguish Algerian/Moroccan darija dialects and understand certain nearly map-disappeared coastal towns, regional cultures, and regional conflicts (a North African dialect roleplay scenario on r).

Review and scope

  • This is a single anonymous user in a single setting (SillyTavern roleplay, non-thinking mode), so the conclusion does not represent all tasks. "Zero censorship" occurred under that mode and prompt setup; comments also discuss censorship differences between the official API (which requires phone-number registration) and OpenRouter, with no consensus.

  • Where it aligns with the official positioning: strong instruction following and good Agent/long-range task performance (official SWE-bench, etc.); where it differs from official promotion: the official source does not mention the friction of needing extremely specific requests for censored content.

  • Task fit: suitable for creative writing, RP, and long-form narratives (add your own anti-repetition instructions); unsuitable for compliance scenarios that require explicit censorship boundaries.

  • Related: another r/SillyTavernAI thread (1u4ts39) reports that during multi-character RP it "likes to insert characters who are not present," which can serve as an additional negative observation for multi-character scenarios.

What this supports

  • This is a single anonymous user in a single setting (SillyTavern roleplay, non-thinking mode), so the conclusion does not represent all tasks. "Zero censorship" occurred under that mode and prompt setup; comments also discuss censorship differences between the official API (which requires phone-number registration) and OpenRouter, with no consensus.
  • Where it aligns with the official positioning: strong instruction following and good Agent/long-range task performance (official SWE-bench, etc.); where it differs from official promotion: the official source does not mention the friction of needing extremely specific requests for censored content.

What this does not support

  • Task fit: suitable for creative writing, RP, and long-form narratives (add your own anti-repetition instructions); unsuitable for compliance scenarios that require explicit censorship boundaries.
  • Related: another r/SillyTavernAI thread (1u4ts39) reports that during multi-character RP it "likes to insert characters who are not present," which can serve as an additional negative observation for multi-character scenarios.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit (r/SillyTavernAI) · u/ (anonymous, no byline on the original post) — "Testing this model for a week" · Original publication date Unknown · Site edit date 2026-09-20

Open original source

LongCat 2.0

Compare LongCat 2.0 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

LongCat 2.0: what changed, where to use it, and what the price misses

LongCat 2.0 combines 1M context, open weights, and low provider pricing with real questions about tooling, data terms, and operational cost.

Related reviews

OpenRouter Channel Data: LongCat-2.0 Pricing, Measured Performance, and Third-Party Benchmarks (Artificial Analysis)The OpenRouter page provides a third-party view beyond the official figures: LongCat-2.0 is listed at $0.30/$1.20 per 1M tokens (with a 60% discount at collection time), while the actual weighted transaction price for input was only $0.03872/M (88.9% cache-hit rate); throughput was P50 29 tok/s, three-day availability 99.93%, and tool-call error rate 0.90%, with real traffic mainly coming from Hermes Agent (7.77B tokens) and Claude Code (3.31B tokens).LongCat-2.0 Official Model Card: Specifications and Official Benchmarks (Including Comparison Tables with Gemini/GPT-5.5/Claude Opus)The official model card is the primary authoritative source for judging LongCat-2.0's suitable tasks: it scores 59.5 on SWE-bench Pro, ahead of GPT-5.5 (58.6) and Gemini 3.1 Pro (54.2), and reaches 70.8 on Terminal-Bench 2.1. However, it trails GPT-5.5 and Claude Opus 4.8 on several benchmarks including BrowseComp, GPQA, and IFEval—in short, it is strong at coding and agent tasks, but not a leader in retrieval and general reasoning.LongCat-2.0 Official Technical Blog: Architecture, Training on Domestic Compute, and Inference Deployment (Release Notes)The official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.eesel Independent Review: LongCat-2.0's Agent Reliability and Hard Blockers to Production DeploymentThis independent review separates LongCat-2.0 into two questions: "can the model complete Agent work?" and "can the product enter enterprise production?" Public user reports support it as an inexpensive, stable coding executor, but its context specifications, tool contract, and data-governance documentation are insufficient to pass a sensitive-data production review.LongCat-2.0 API Platform Quick Start (Official Quick Start + Chat Completions Reference + Pricing)The LongCat Claude Code guide configures a compatible endpoint and keeps the first task in a disposable worktree.LongCat-2.0 Chat Template and Tool-Calling Configuration (Official Hugging Face Model Card)The official model card’s chat template and tool-call examples are converted into a local inference configuration check.Claude Code Integration with LongCat-2.0 (Official Documentation)The official LongCat integration guide configures a named client and keeps the first run observable and reversible.Official Account Showcase: Five “One-Prompt Generation” Creative Projects (Voxel/3D/CG/Landing Page/Mini-game)Source “Official Account Showcase: Five “One-Prompt Generation” Creative Projects (Voxel/3D/CG/Landing Page/Mini-game)” is organized as an executable task guide; its environment, inputs, and acceptance boundary follow the source.