Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

LongCat 2.0 · Community source · Personal experience

r/hermesagent PSA: LongCat 2.0 Reasoning-Tier Bug and Model Positioning (Between DeepSeek V4 Flash/Pro)

Testing confirmed that LongCat-2.0's API accepts only three reasoning-effort tiers, `low/med/high`. When it receives another tier (such as `xhigh` from the DeepSeek family), it returns a malformed 200 response instead of an error, causing Hermes Agent to silently fall back to a fallback provider. The same user positioned its capabilities between DeepSeek-V4-Flash and DeepSeek-V4-Pro, with good caching and cheap PAYG pricing.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model/version
LongCat-2.0; source date: 2026-08-18.
Harness/task
Model positioning (user's subjective view, based on actual use); "After using it for a while, my feeling is that it's (capability-wise) between Deepseek-4-Flash and Deepseek-4-Pro, and the pricing is too."
Sample/gaps
Limitations noted: The capability positioning (between DS4-Flash and DS4-Pro) is subjective and complements AlphaSignal's claim that DeepSeek V4-Pro is cheaper (Review 04) and the OpenRouter third-party index (Review 03); these can be cross-checked against one another.; Task fit: suitable for cache-friendly Agent loops and PAYG cost-sensitive scenarios; Hermes users who switch models frequently should watch for the configuration pitfall above.

Key data and applicable tasks

One-sentence takeaway

Testing confirmed that LongCat-2.0's API accepts only three reasoning-effort tiers, low/med/high. When it receives another tier (such as xhigh from the DeepSeek family), it returns a malformed 200 response instead of an error, causing Hermes Agent to silently fall back to a fallback provider. The same user positioned its capabilities between DeepSeek-V4-Flash and DeepSeek-V4-Pro, with good caching and cheap PAYG pricing.

Original report (key points from the original)

Model positioning (user's subjective view, based on actual use)

  • "After using it for a while, my feeling is that it's (capability-wise) between Deepseek-4-Flash and Deepseek-4-Pro, and the pricing is too."

  • They like that it offers a PAYG (pay-as-you-go) plan; caching works as well as Deepseek's, and the API is "extremely cheap" in actual use.

  • "Definitely going to be in my regular model rotation."

Bug details (core of the PSA)

  • Scenario: the Agent was switched from DS4P to LongCat-2.0 and immediately hit the fallback provider.

  • Error (agent.log, excerpt from the original):

agent.conversation_loop: API call failed (attempt 1/3) error_type=RuntimeError
provider=custom base_url=https://api.longcat.chat/openai/v1 model=LongCat-2.0
summary=Provider returned an empty stream with no finish_reason
(possible upstream error or malformed SSE response).
  • Root cause: LongCat accepts only low, med, high reasoning levels; the user had previously set xhigh on DeepSeek-4-Pro to obtain maximum thinking. When Hermes switched models, it did not change the reasoning effort, so it sent xhigh to LongCat; LongCat returned a malformed 200 response (an empty stream with no finish_reason) instead of an error response.

  • The issue was reported to LongCat, which acknowledged it ("They appear to have acknowledged the issue").

  • Lesson: when frequently switching models in Hermes and different models use different reasoning-effort tiers, change the effort back to low/med/high before switching to LongCat.

Review and scope

  • This is a single-point report from an anonymous user, with the error log quoted from the original. It differs from the thinking: {"type":"enabled"/"disabled"} field in the official API reference (Prompt Directory 01) — this report concerns the reasoning-effort concept on the Hermes side, and the official documentation does not document how the two map to each other. Take care when switching models.

  • The capability positioning (between DS4-Flash and DS4-Pro) is subjective and complements AlphaSignal's claim that DeepSeek V4-Pro is cheaper (Review 04) and the OpenRouter third-party index (Review 03); these can be cross-checked against one another.

  • Task fit: suitable for cache-friendly Agent loops and PAYG cost-sensitive scenarios; Hermes users who switch models frequently should watch for the configuration pitfall above.

What this supports

  • This is a single-point report from an anonymous user, with the error log quoted from the original. It differs from the `thinking: {"type":"enabled"/"disabled"}` field in the official API reference (Prompt Directory 01) — this report concerns the reasoning-effort concept on the Hermes side, and the official documentation does not document how the two map to each other. Take care when switching models.
  • The capability positioning (between DS4-Flash and DS4-Pro) is subjective and complements AlphaSignal's claim that DeepSeek V4-Pro is cheaper (Review 04) and the OpenRouter third-party index (Review 03); these can be cross-checked against one another.

What this does not support

  • The capability positioning (between DS4-Flash and DS4-Pro) is subjective and complements AlphaSignal's claim that DeepSeek V4-Pro is cheaper (Review 04) and the OpenRouter third-party index (Review 03); these can be cross-checked against one another.
  • Task fit: suitable for cache-friendly Agent loops and PAYG cost-sensitive scenarios; Hermes users who switch models frequently should watch for the configuration pitfall above.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit (r/hermesagent) · u/ (anonymous original poster, PSA post) · Original publication date Unknown · Site edit date 2026-09-20

Open original source

LongCat 2.0

Compare LongCat 2.0 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

LongCat 2.0: what changed, where to use it, and what the price misses

LongCat 2.0 combines 1M context, open weights, and low provider pricing with real questions about tooling, data terms, and operational cost.

Related reviews

OpenRouter Channel Data: LongCat-2.0 Pricing, Measured Performance, and Third-Party Benchmarks (Artificial Analysis)The OpenRouter page provides a third-party view beyond the official figures: LongCat-2.0 is listed at $0.30/$1.20 per 1M tokens (with a 60% discount at collection time), while the actual weighted transaction price for input was only $0.03872/M (88.9% cache-hit rate); throughput was P50 29 tok/s, three-day availability 99.93%, and tool-call error rate 0.90%, with real traffic mainly coming from Hermes Agent (7.77B tokens) and Claude Code (3.31B tokens).LongCat-2.0 Official Model Card: Specifications and Official Benchmarks (Including Comparison Tables with Gemini/GPT-5.5/Claude Opus)The official model card is the primary authoritative source for judging LongCat-2.0's suitable tasks: it scores 59.5 on SWE-bench Pro, ahead of GPT-5.5 (58.6) and Gemini 3.1 Pro (54.2), and reaches 70.8 on Terminal-Bench 2.1. However, it trails GPT-5.5 and Claude Opus 4.8 on several benchmarks including BrowseComp, GPQA, and IFEval—in short, it is strong at coding and agent tasks, but not a leader in retrieval and general reasoning.LongCat-2.0 Official Technical Blog: Architecture, Training on Domestic Compute, and Inference Deployment (Release Notes)The official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.eesel Independent Review: LongCat-2.0's Agent Reliability and Hard Blockers to Production DeploymentThis independent review separates LongCat-2.0 into two questions: "can the model complete Agent work?" and "can the product enter enterprise production?" Public user reports support it as an inexpensive, stable coding executor, but its context specifications, tool contract, and data-governance documentation are insufficient to pass a sensitive-data production review.LongCat-2.0 API Platform Quick Start (Official Quick Start + Chat Completions Reference + Pricing)The LongCat Claude Code guide configures a compatible endpoint and keeps the first task in a disposable worktree.LongCat-2.0 Chat Template and Tool-Calling Configuration (Official Hugging Face Model Card)The official model card’s chat template and tool-call examples are converted into a local inference configuration check.Claude Code Integration with LongCat-2.0 (Official Documentation)The official LongCat integration guide configures a named client and keeps the first run observable and reversible.Official Account Showcase: Five “One-Prompt Generation” Creative Projects (Voxel/3D/CG/Landing Page/Mini-game)Source “Official Account Showcase: Five “One-Prompt Generation” Creative Projects (Voxel/3D/CG/Landing Page/Mini-game)” is organized as an executable task guide; its environment, inputs, and acceptance boundary follow the source.