The OpenRouter page provides a third-party view beyond the official figures: LongCat-2.0 is listed at $0.30/$1.20 per 1M tokens (with a 60% discount at collection time), while the actual weighted transaction price for input was only $0.03872/M (88.9% cache-hit rate); throughput was P50 29 tok/s, three-day availability 99.93%, and tool-call error rate 0.90%, with real traffic mainly coming from Hermes Agent (7.77B tokens) and Claude Code (3.31B tokens).
OpenRouter (third-party model routing platform) · Read evidenceLongCat 2.0 · Reviews and evidence
Which LongCat 2.0 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
The official model card is the primary authoritative source for judging LongCat-2.0's suitable tasks: it scores 59.5 on SWE-bench Pro, ahead of GPT-5.5 (58.6) and Gemini 3.1 Pro (54.2), and reaches 70.8 on Terminal-Bench 2.1. However, it trails GPT-5.5 and Claude Opus 4.8 on several benchmarks including BrowseComp, GPQA, and IFEval—in short, it is strong at coding and agent tasks, but not a leader in retrieval and general reasoning.
Hugging Face (meituan-longcat/LongCat-2.0) · Read evidenceThe official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.
LongCat official blog (longcat.chat) · Read evidenceFull reviews and related reading
Selected evidence
OpenRouter Channel Data: LongCat-2.0 Pricing, Measured Performance, and Third-Party Benchmarks (Artificial Analysis)
The OpenRouter page provides a third-party view beyond the official figures: LongCat-2.0 is listed at $0.30/$1.20 per 1M tokens (with a 60% discount at collection time), while the actual weighted transaction price for input was only $0.03872/M (88.9% cache-hit rate); throughput was P50 29 tok/s, three-day availability 99.93%, and tool-call error rate 0.90%, with real traffic mainly coming from Hermes Agent (7.77B tokens) and Claude Code (3.31B tokens).
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-07-20.
- Harness/task
- Model and pricing; Model ID: `meituan/longcat-2.0` (page version 20260720); 48B active / 1.6T total parameters MoE; 1M context; text only.
- Sample/gaps
- Limitations noted: Artificial Analysis's Coding Index 45.3 (better than 49% of models) is clearly below the impression created by the official SWE-bench Pro score of 59.5. The benchmark sets differ, and the third-party index does not rank LongCat particularly highly, which is an important correction when judging "which tasks it suits."; A single provider (AtlasCloud) means there is no multi-route redundancy; the 99.93% availability is a snapshot covering roughly the past three days.
LongCat-2.0 Official Model Card: Specifications and Official Benchmarks (Including Comparison Tables with Gemini/GPT-5.5/Claude Opus)
The official model card is the primary authoritative source for judging LongCat-2.0's suitable tasks: it scores 59.5 on SWE-bench Pro, ahead of GPT-5.5 (58.6) and Gemini 3.1 Pro (54.2), and reaches 70.8 on Terminal-Bench 2.1. However, it trails GPT-5.5 and Claude Opus 4.8 on several benchmarks including BrowseComp, GPQA, and IFEval—in short, it is strong at coding and agent tasks, but not a leader in retrieval and general reasoning.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-06-30.
- Harness/task
- Architecture: MoE, 1.6T total parameters, about 48B active per token (dynamic 33B–56B, according to the official X account); HF metadata lists a 1.8T model size (including 135B N-gram Embedding parameters, BF16).; Context: Native 1M tokens (trained on hundreds of billions of tokens of million-token-context data).
- Sample/gaps
- Limitations noted: Areas behind: FORTE/RWSearch/BrowseComp all trail GPT-5.5 (77.8/85.3/84.4 vs. 73.2/78.8/79.9); GPQA-diamond and IFEval trail GPT-5.5 and Gemini 3.1 Pro; Claude Opus 4.8 leads LongCat by nearly 10 points on SWE-bench Pro (69.2).; Overall: In the official chart, the only model LongCat-2.0 consistently beats is Gemini 3.1 Pro (AlphaSignal's conclusion); against GPT-5.5 it narrowly wins only on SWE-bench Pro.
LongCat-2.0 Official Technical Blog: Architecture, Training on Domestic Compute, and Inference Deployment (Release Notes)
The official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-06-30.
- Harness/task
- Release claims (consistent with the model card); A 1.6T-total-parameter MoE with approximately 48B parameters activated per token; all full training and large-scale deployment use domestic compute clusters.
- Sample/gaps
- Limitations noted: Scope: the blog is an engineering note rather than an independent evaluation. Official throughput/reliability figures (such as a 70%+ reduction in failure rate) have no third-party verification; the deployment approach targets very large clusters and offers individual developers mainly the SGLang cookbook as a reference.; Related documents: Review 01 (official benchmark table), Review 04 (AlphaSignal verification).
eesel Independent Review: LongCat-2.0's Agent Reliability and Hard Blockers to Production Deployment
This independent review separates LongCat-2.0 into two questions: "can the model complete Agent work?" and "can the product enter enterprise production?" Public user reports support it as an inexpensive, stable coding executor, but its context specifications, tool contract, and data-governance documentation are insufficient to pass a sensitive-data production review.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-08-04.
- Harness/task
- This independent review separates LongCat-2.0 into two questions: "can the model complete Agent work?" and "can the product enter enterprise production?" Public user reports support it as an inexpensive, stable coding executor, but its context specifications, tool contract, and data-governance documentation are insufficient to pass a sensitive-data production review.
- Sample/gaps
- Limitations noted: The author works at eesel AI, and the article contains promotion for its product. The data-governance checks can be reviewed, but purchase recommendations should be considered separately from the provider's original policies and the opinions of legal and security teams.; When reproducing, save: the official config.json, API response headers and error bodies, harness configuration, input token count, tool schema, provider policy pages, and collection date.
All sources
All sources
OpenRouter Channel Data: LongCat-2.0 Pricing, Measured Performance, and Third-Party Benchmarks (Artificial Analysis)
The OpenRouter page provides a third-party view beyond the official figures: LongCat-2.0 is listed at $0.30/$1.20 per 1M tokens (with a 60% discount at collection time), while the actual weighted transaction price for input was only $0.03872/M (88.9% cache-hit rate); throughput was P50 29 tok/s, three-day availability 99.93%, and tool-call error rate 0.90%, with real traffic mainly coming from Hermes Agent (7.77B tokens) and Claude Code (3.31B tokens).
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-07-20.
- Harness/task
- Model and pricing; Model ID: `meituan/longcat-2.0` (page version 20260720); 48B active / 1.6T total parameters MoE; 1M context; text only.
- Sample/gaps
- Limitations noted: Artificial Analysis's Coding Index 45.3 (better than 49% of models) is clearly below the impression created by the official SWE-bench Pro score of 59.5. The benchmark sets differ, and the third-party index does not rank LongCat particularly highly, which is an important correction when judging "which tasks it suits."; A single provider (AtlasCloud) means there is no multi-route redundancy; the 99.93% availability is a snapshot covering roughly the past three days.
LongCat-2.0 Official Model Card: Specifications and Official Benchmarks (Including Comparison Tables with Gemini/GPT-5.5/Claude Opus)
The official model card is the primary authoritative source for judging LongCat-2.0's suitable tasks: it scores 59.5 on SWE-bench Pro, ahead of GPT-5.5 (58.6) and Gemini 3.1 Pro (54.2), and reaches 70.8 on Terminal-Bench 2.1. However, it trails GPT-5.5 and Claude Opus 4.8 on several benchmarks including BrowseComp, GPQA, and IFEval—in short, it is strong at coding and agent tasks, but not a leader in retrieval and general reasoning.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-06-30.
- Harness/task
- Architecture: MoE, 1.6T total parameters, about 48B active per token (dynamic 33B–56B, according to the official X account); HF metadata lists a 1.8T model size (including 135B N-gram Embedding parameters, BF16).; Context: Native 1M tokens (trained on hundreds of billions of tokens of million-token-context data).
- Sample/gaps
- Limitations noted: Areas behind: FORTE/RWSearch/BrowseComp all trail GPT-5.5 (77.8/85.3/84.4 vs. 73.2/78.8/79.9); GPQA-diamond and IFEval trail GPT-5.5 and Gemini 3.1 Pro; Claude Opus 4.8 leads LongCat by nearly 10 points on SWE-bench Pro (69.2).; Overall: In the official chart, the only model LongCat-2.0 consistently beats is Gemini 3.1 Pro (AlphaSignal's conclusion); against GPT-5.5 it narrowly wins only on SWE-bench Pro.
LongCat-2.0 Official Technical Blog: Architecture, Training on Domestic Compute, and Inference Deployment (Release Notes)
The official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-06-30.
- Harness/task
- Release claims (consistent with the model card); A 1.6T-total-parameter MoE with approximately 48B parameters activated per token; all full training and large-scale deployment use domestic compute clusters.
- Sample/gaps
- Limitations noted: Scope: the blog is an engineering note rather than an independent evaluation. Official throughput/reliability figures (such as a 70%+ reduction in failure rate) have no third-party verification; the deployment approach targets very large clusters and offers individual developers mainly the SGLang cookbook as a reference.; Related documents: Review 01 (official benchmark table), Review 04 (AlphaSignal verification).
eesel Independent Review: LongCat-2.0's Agent Reliability and Hard Blockers to Production Deployment
This independent review separates LongCat-2.0 into two questions: "can the model complete Agent work?" and "can the product enter enterprise production?" Public user reports support it as an inexpensive, stable coding executor, but its context specifications, tool contract, and data-governance documentation are insufficient to pass a sensitive-data production review.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-08-04.
- Harness/task
- This independent review separates LongCat-2.0 into two questions: "can the model complete Agent work?" and "can the product enter enterprise production?" Public user reports support it as an inexpensive, stable coding executor, but its context specifications, tool contract, and data-governance documentation are insufficient to pass a sensitive-data production review.
- Sample/gaps
- Limitations noted: The author works at eesel AI, and the article contains promotion for its product. The data-governance checks can be reviewed, but purchase recommendations should be considered separately from the provider's original policies and the opinions of legal and security teams.; When reproducing, save: the official config.json, API response headers and error bodies, harness configuration, input token count, tool schema, provider policy pages, and collection date.
AlphaSignal Deep Dive: Owl Alpha's True Identity and a Reality Check on LongCat-2.0's Official Claims
This deep-dive review, published the day after launch, establishes the most important background fact — the anonymous free model "Owl Alpha," which ran on OpenRouter for two months, was LongCat-2.0 (processing approximately 10 trillion tokens per month and ranking first on the Hermes Agent leaderboard) — while breaking down the official claims one by one: the narrow SWE-bench Pro win over GPT-5.5 was an official self-test, the weights were still only "planned" when the review was published, and the open-source fiel
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-07-01.
- Harness/task
- This deep-dive review, published the day after launch, establishes the most important background fact — the anonymous free model "Owl Alpha," which ran on OpenRouter for two months, was LongCat-2.0 (processing approximately 10 trillion tokens per month and ranking first on the Hermes Agent leaderboard) — while breaking down the official claims one by one: the narrow SWE-bench Pro win over GPT-5.5 was an official self-test, the weights were still only "planned" when the review was published, and the open-source fiel
- Sample/gaps
- Limitations noted: The article is analytical commentary rather than a controlled evaluation. Its figures are relayed from official and platform sources, but key data such as "10 trillion tokens/month" and "GLM-5.2 62.1" have no independent original-source links (the author says source links are in the first reply); treat them as second-hand when citing.; It is best read alongside the official documents (Reviews 01 and 02): the official documents provide the full details, while this article provides a critical framework.
r/SillyTavernAI Field Test: One Week of LongCat 2.0 Roleplay (Writing/Jailbreaks/Repetition Tendency)
A one-week roleplay test found that LongCat 2.0 was "a jackpot" for creative writing: faithful instruction following, coherent stories, no hard refusals, and dry, non-sensational narration. It also had two clear flaws — excessive fidelity to supplied information leading to repetitive inertia, and the need for extremely specific requests to reach jailbreak content (the author had to use an exceptionally strong jailbreak prompt).
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-08-18.
- Harness/task
- Pros; Faithful and accurate instruction following: it can gracefully handle very complex worldbuilding setups.
- Sample/gaps
- Limitations noted: Task fit: suitable for creative writing, RP, and long-form narratives (add your own anti-repetition instructions); unsuitable for compliance scenarios that require explicit censorship boundaries.; Related: another r/SillyTavernAI thread (1u4ts39) reports that during multi-character RP it "likes to insert characters who are not present," which can serve as an additional negative observation for multi-character scenarios.
r/hermesagent PSA: LongCat 2.0 Reasoning-Tier Bug and Model Positioning (Between DeepSeek V4 Flash/Pro)
Testing confirmed that LongCat-2.0's API accepts only three reasoning-effort tiers, `low/med/high`. When it receives another tier (such as `xhigh` from the DeepSeek family), it returns a malformed 200 response instead of an error, causing Hermes Agent to silently fall back to a fallback provider. The same user positioned its capabilities between DeepSeek-V4-Flash and DeepSeek-V4-Pro, with good caching and cheap PAYG pricing.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-08-18.
- Harness/task
- Model positioning (user's subjective view, based on actual use); "After using it for a while, my feeling is that it's (capability-wise) between Deepseek-4-Flash and Deepseek-4-Pro, and the pricing is too."
- Sample/gaps
- Limitations noted: The capability positioning (between DS4-Flash and DS4-Pro) is subjective and complements AlphaSignal's claim that DeepSeek V4-Pro is cheaper (Review 04) and the OpenRouter third-party index (Review 03); these can be cross-checked against one another.; Task fit: suitable for cache-friendly Agent loops and PAYG cost-sensitive scenarios; Hermes users who switch models frequently should watch for the configuration pitfall above.
r/vibecoding field test: LongCat 2.0's "insane" pricing — a real bill with free cache hits
A user tested a $2 package containing 50 million tokens: a single task used 27 million prompt tokens, but only 570,000 tokens were deducted because of cache hits. The official rule — "cache hits are not billed; only misses and output count" — expands effective usage to roughly 2 billion tokens in a real agent workflow, making this the most counterintuitive part of LongCat's pricing.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-08-18.
- Harness/task
- Motivation: The user was intrigued by Meituan's LongCat 2.0 but hesitant because there were not many benchmarks. Everyone was praising Owl Alpha, though, and "it turned out to be the same model."; Purchase: The official 50M-token package costs $2 (the actual charge was €0.86, "not sure why"); installing AliPay on a phone was required and took about 15 minutes.
- Sample/gaps
- Limitations noted: A user tested a $2 package containing 50 million tokens: a single task used 27 million prompt tokens, but only 570,000 tokens were deducted because of cache hits. The official rule — "cache hits are not billed; only misses and output count" — expands effective usage to roughly 2 billion tokens in a real agent workflow, making this the most counterintuitive part of LongCat's pricing.
X field test: DeepSeek V4 Flash / V4 Pro / LongCat 2.0 in the same physics-and-coding task
The developer put DeepSeek V4 Flash 0731, DeepSeek V4 Pro, and LongCat 2.0 through the same "physics + coding" test (same task, same conditions). The result: "LongCat 2.0 surprised me the most" — it performed far more successfully than expected, and this was his first time encountering it on a free tier.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-08-07.
- Harness/task
- > I put DeepSeek V4 Flash 0731, DeepSeek V4 Pro, and LongCat 2.0 through the same physics + coding test.; > All three models received the same task under the same conditions.
- Sample/gaps
- Limitations noted: The developer put DeepSeek V4 Flash 0731, DeepSeek V4 Pro, and LongCat 2.0 through the same "physics + coding" test (same task, same conditions). The result: "LongCat 2.0 surprised me the most" — it performed far more successfully than expected, and this was his first time encountering it on a free tier.
X (atomic.chat) comparison: LongCat 2.0 vs GPT-5.5 in the same agentic game-development task (Duck Hunt)
In Kilo Code CLI, LongCat 2.0 (open weights, running locally/free) and GPT-5.5 (paid cloud) performed the same task: three agent iterations turned `game.html` into a retro Duck Hunt game with duck waves, ammunition, and physics. Both outputs were "smooth, with no clipping issues, and synchronized graphics/physics/game logic"; the only difference was the bill: $0 for LongCat versus $0.65 for GPT-5.5.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-07-07.
- Harness/task
- In Kilo Code CLI, LongCat 2.0 (open weights, running locally/free) and GPT-5.5 (paid cloud) performed the same task: three agent iterations turned `game.html` into a retro Duck Hunt game with duck waves, ammunition, and physics. Both outputs were "smooth, with no clipping issues, and synchronized graphics/physics/game logic"; the only difference was the bill: $0 for LongCat versus $0.65 for GPT-5.5.
- Sample/gaps
- Limitations noted: In Kilo Code CLI, LongCat 2.0 (open weights, running locally/free) and GPT-5.5 (paid cloud) performed the same task: three agent iterations turned `game.html` into a retro Duck Hunt game with duck waves, ammunition, and physics. Both outputs were "smooth, with no clipping issues, and synchronized graphics/physics/game logic"; the only difference was the bill: $0 for LongCat versus $0.65 for GPT-5.5.
AI Profit Boardroom field test: LongCat 2.0 game-building test and same-task comparison with GLM 5.2
The author's test reached a conclusion opposite to most community sentiment: LongCat 2.0's games were "playable but rough and buggy" (one build even showed a completely black screen), while GLM 5.2's outputs on the same tasks were "cleaner, smoother, and more polished." His recommendation was "worth playing with, not worth switching to" — a negative independent sample that should be read alongside positive evidence about LongCat 2.0.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-05-29.
- Harness/task
- Test method; LongCat 2.0 was asked to build several game demos: Dragon Realm, a Skyrim-style open world, and VoxelCraft. The conclusion was "playable, but buggy and rough — one build even went completely black at certain points."
- Sample/gaps
- Limitations noted: The author's test reached a conclusion opposite to most community sentiment: LongCat 2.0's games were "playable but rough and buggy" (one build even showed a completely black screen), while GLM 5.2's outputs on the same tasks were "cleaner, smoother, and more polished." His recommendation was "worth playing with, not worth switching to" — a negative independent sample that should be read alongside positive evidence about LongCat 2.0.
BenchLM comparison page: GPT-5.5 vs LongCat-2.0 — a boundary note on "no shared benchmarks, no quality verdict"
As of 2026-08-17, BenchLM found no shared third-party benchmark results between GPT-5.5 and LongCat-2.0 (38 for GPT-5.5 and 0 for LongCat-2.0), so "the public evidence does not support any quality verdict." This is an authoritative boundary reminder for claims that "LongCat 2.0 beats GPT-5.5": the official SWE-bench Pro result of 59.5 > 58.6 has not yet been retested by any independent source.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-08-17.
- Harness/task
- As of 2026-08-17, BenchLM found no shared third-party benchmark results between GPT-5.5 and LongCat-2.0 (38 for GPT-5.5 and 0 for LongCat-2.0), so "the public evidence does not support any quality verdict." This is an authoritative boundary reminder for claims that "LongCat 2.0 beats GPT-5.5": the official SWE-bench Pro result of 59.5 > 58.6 has not yet been retested by any independent source.
- Sample/gaps
- Limitations noted: As of 2026-08-17, BenchLM found no shared third-party benchmark results between GPT-5.5 and LongCat-2.0 (38 for GPT-5.5 and 0 for LongCat-2.0), so "the public evidence does not support any quality verdict." This is an authoritative boundary reminder for claims that "LongCat 2.0 beats GPT-5.5": the official SWE-bench Pro result of 59.5 > 58.6 has not yet been retested by any independent source.
r/AIToolsPerformance: Pricing Discussion Comparing LongCat 2.0 with Kimi K3 and Other 1M-Context Models
On the day of its release, the community noticed that with the same 1M context, LongCat 2.0 ($0.30/$1.20) was about 10 times cheaper than Kimi K3 ($3.00/$15.00) for input and about 12 times cheaper for output. It was then "the cheapest 1M+ context model" on OpenRouter—but the post also warned that the listing page showed no quality data: "aggressive pricing could be a land-grab, or it could be genuinely efficient."
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-07-20.
- Harness/task
- On the day of its release, the community noticed that with the same 1M context, LongCat 2.0 ($0.30/$1.20) was about 10 times cheaper than Kimi K3 ($3.00/$15.00) for input and about 12 times cheaper for output. It was then "the cheapest 1M+ context model" on OpenRouter—but the post also warned that the listing page showed no quality data: "aggressive pricing could be a land-grab, or it could be genuinely efficient."
- Sample/gaps
- Limitations noted: On the day of its release, the community noticed that with the same 1M context, LongCat 2.0 ($0.30/$1.20) was about 10 times cheaper than Kimi K3 ($3.00/$15.00) for input and about 12 times cheaper for output. It was then "the cheapest 1M+ context model" on OpenRouter—but the post also warned that the listing page showed no quality data: "aggressive pricing could be a land-grab, or it could be genuinely efficient."
r/LocalLLaMA Discussion: Weight Releases, Download Size, and Speculation About Domestic "AI ASIC Superpods"
From the time of release, r/LocalLLaMA focused on two issues: weights and quantization (3.55 TB for the full BF16 model, 2.05 TB for FP8, with official INT8/FP8 quantized versions) and whose chips power the "AI ASIC superpods" (the community inferred Huawei Ascend 910C from the Huawei HCCL acknowledgment and the term "superpod"). These are important community signals for judging whether LongCat-2.0 can be deployed in practice.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-07.
- Harness/task
- From the time of release, r/LocalLLaMA focused on two issues: weights and quantization (3.55 TB for the full BF16 model, 2.05 TB for FP8, with official INT8/FP8 quantized versions) and whose chips power the "AI ASIC superpods" (the community inferred Huawei Ascend 910C from the Huawei HCCL acknowledgment and the term "superpod"). These are important community signals for judging whether LongCat-2.0 can be deployed in practice.
- Sample/gaps
- Limitations noted: From the time of release, r/LocalLLaMA focused on two issues: weights and quantization (3.55 TB for the full BF16 model, 2.05 TB for FP8, with official INT8/FP8 quantized versions) and whose chips power the "AI ASIC superpods" (the community inferred Huawei Ascend 910C from the Huawei HCCL acknowledgment and the term "superpod"). These are important community signals for judging whether LongCat-2.0 can be deployed in practice.
Hacker News Single-Question Comparison: LongCat-2.0's Scientific-Reasoning Error and the Boundaries of Test Design
Hacker News' single-question comparison offers a reproducible but non-ranking warning sample: on a nuclear-fuel-selection question, LongCat-2.0 gave reasons the author judged incorrect, while Qwen 3.7 Plus and Gemini Flash gave different answers. The comments pointed out that the question's semantics, factual background, and n=1 design were all insufficient to support a general conclusion.
Unverified: the original source could not be rechecked.
- Model/version
- LongCat-2.0; source date: 2026-06-29.
- Harness/task
- The original author presented U-235 and Pu-241 (both mixed with 95% U-238) as two possible nuclear-reactor fuels and asked which one should be chosen and why; the original question can be checked directly on the page.; The author described LongCat-2.0's answer as "very well-written reasoning but an incorrect conclusion"; it chose Pu-241.
- Sample/gaps
- Limitations noted: Hacker News' single-question comparison offers a reproducible but non-ranking warning sample: on a nuclear-fuel-selection question, LongCat-2.0 gave reasons the author judged incorrect, while Qwen 3.7 Plus and Gemini Flash gave different answers. The comments pointed out that the question's semantics, factual background, and n=1 design were all insufficient to support a general conclusion.
LongCat 2.0
Compare LongCat 2.0 in Tabbit
Model access, features, and permissions depend on your current client account.