OpenRouter (third-party model routing platform)Platform telemetry
OpenRouter Channel Data: LongCat-2.0 Pricing, Measured Performance, and Third-Party Benchmarks (Artificial Analysis)
The OpenRouter page provides a third-party view beyond the official figures: LongCat-2.0 is listed at $0.30/$1.20 per 1M tokens (with a 60% discount at collection time), while the actual weighted transaction price for input was only $0.03872/M (88.9% cache-hit rate); throughput was P50 29 tok/s, three-day availability 99.93%, and tool-call error rate 0.90%, with real traffic mainly coming from Hermes Agent (7.77B tokens) and Claude Code (3.31B tokens).
- Evidence
- Platform telemetry
- Boundary
- Artificial Analysis's Coding Index 45.3 (better than 49% of models) is clearly below the impression created by the official SWE-bench Pro score of 59.5. The benchmark sets differ, and the third-party index does not rank LongCat particularly highly, which is an important correction when judging "which tasks it suits."
Hugging Face (meituan-longcat/LongCat-2.0)Vendor report
LongCat-2.0 Official Model Card: Specifications and Official Benchmarks (Including Comparison Tables with Gemini/GPT-5.5/Claude Opus)
The official model card is the primary authoritative source for judging LongCat-2.0's suitable tasks: it scores 59.5 on SWE-bench Pro, ahead of GPT-5.5 (58.6) and Gemini 3.1 Pro (54.2), and reaches 70.8 on Terminal-Bench 2.1. However, it trails GPT-5.5 and Claude Opus 4.8 on several benchmarks including BrowseComp, GPQA, and IFEval—in short, it is strong at coding and agent tasks, but not a leader in retrieval and general reasoning.
- Evidence
- Vendor report
- Boundary
- Areas behind: FORTE/RWSearch/BrowseComp all trail GPT-5.5 (77.8/85.3/84.4 vs. 73.2/78.8/79.9); GPQA-diamond and IFEval trail GPT-5.5 and Gemini 3.1 Pro; Claude Opus 4.8 leads LongCat by nearly 10 points on SWE-bench Pro (69.2).
LongCat official blog (longcat.chat)Vendor report
LongCat-2.0 Official Technical Blog: Architecture, Training on Domestic Compute, and Inference Deployment (Release Notes)
The official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.
- Evidence
- Vendor report
- Boundary
- Scope: the blog is an engineering note rather than an independent evaluation. Official throughput/reliability figures (such as a 70%+ reduction in failure rate) have no third-party verification; the deployment approach targets very large clusters and offers individual developers mainly the SGLang cookbook as a reference.