Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

LongCat 2.0 · Media / benchmark · Vendor report

LongCat-2.0 Official Technical Blog: Architecture, Training on Domestic Compute, and Inference Deployment (Release Notes)

The official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkVendor reportEdited 2026-09-20

Test conditions

Model/version
LongCat-2.0; source date: 2026-06-30.
Harness/task
Release claims (consistent with the model card); A 1.6T-total-parameter MoE with approximately 48B parameters activated per token; all full training and large-scale deployment use domestic compute clusters.
Sample/gaps
Limitations noted: Scope: the blog is an engineering note rather than an independent evaluation. Official throughput/reliability figures (such as a 70%+ reduction in failure rate) have no third-party verification; the deployment approach targets very large clusters and offers individual developers mainly the SGLang cookbook as a reference.; Related documents: Review 01 (official benchmark table), Review 04 (AlphaSignal verification).

Key data and applicable tasks

One-sentence takeaway

The official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.

Official core content

Release claims (consistent with the model card)

  • A 1.6T-total-parameter MoE with approximately 48B parameters activated per token; all full training and large-scale deployment use domestic compute clusters.

  • Pretraining: more than 50,000 domestic compute chips, taking more than a month, on 35T+ tokens, with no rollback throughout and no unrecoverable loss spikes.

  • Long-range capabilities: LongCat sparse attention + training on hundreds of billions of tokens of million-token-context data + dedicated post-training → strong performance on coding and agent tasks.

  • Deeply adapted to Claude Code, OpenClaw, and Hermes; online experience at https://longcat.chat; API access at https://longcat.chat/platform/docs/.

Architecture upgrades (key points from the official text)

  1. LongCat Sparse Attention (LSA) — evolved from DeepSeek Sparse Attention (DSA), targeting the bottlenecks in DSA's Lightning Indexer: "discontinuous index outputs + quadratic index scoring":

    • Streaming-aware Indexing (SI): hardware-aligned contiguous access + dynamic random selection; turns fragmented GPU-memory access into sequential reads and merges HBM accesses.

    • Cross-Layer Indexing (CLI): adjacent attention layers have substantially consistent distributions of salient tokens, so one index can be reused across multiple consecutive layers (supported by cross-layer distillation during training).

    • Hierarchical Indexing (HI): two-stage coarse-to-fine scoring (block-level coarse retrieval → fine selection within candidates), enabled on demand.

    • The three components are orthogonal and can be toggled independently; all are extended to 3-step MTP speculative decoding (the target model shares one index every 2 layers; the 3 draft steps share one index).

  2. N-gram Embedding: inherited from LongCat-Flash-Lite; n-gram size=5, 135B parameters; expands the embedding space by more than 100x. Expansion principles: MoE sparsity has passed its sweet spot (~97%), while the N-gram share is constrained to the optimal range (<10% of total parameters; experiments show the advantage disappears above 50%).

Training (domestic compute infrastructure)

  • 6D parallelism: EMBP parallelism for N-gram Embedding is added alongside TP/CP/EP/DP/PP.

  • Supernodes: physical supernodes contain up to 48 machines, fully interconnected within a node and connected between nodes via RoCE; this brings approximately a 30% pretraining-throughput improvement.

  • Memory optimization: ZeRO-1, selective recomputation, automatic OOM offloading, and routing padding tokens to zero-computation experts.

  • Large-scale deployment of the Muon optimizer (TP parallelism, deduplication of DP state, and dedicated optimization of symmetric matrix-multiplication kernels).

  • Long context: LSA warm-up forward-only + KL loss; all-gather context parallelism scaled beyond 512 ways; overlapping computation and communication (ScMoE, and overlap between top-k indexing and KV all-gather).

  • Reliability: deterministic operators with dual communication/computation paths; segmented binary-tree accumulation for reduction operators; bit-flip detection; automatic fault identification, traffic switching, and recovery, with repaired links reused after stress testing.

Inference and deployment (service layer)

  • Million-token-context inference under memory constraints: absorb computation mode (prefill/decode), parallel indexer and MLA prolog, and multi-GPU KV-cache partitioning (KVP); ScMoE uses core-control capabilities to fully parallelize dense and MoE streams; super kernel reduces operator-launch overhead; Weight Prefetch uses a large L2 cache to hide I/O latency; 200Gbps network cards transfer layer-wise KV caches.

  • Deployment: prefill–decode (PD) disaggregation; Prefill nodes use multi-node Chunked Pipeline Parallel (CPP) to reduce the Expert-Parallel domain, plus Attention Sequence Parallelism (SP); Decode nodes use KVP to shard the KV cache.

  • Official deployment entry points: SGLang cookbook for GPUs, SGLang-FluentLLM for NPUs; weights released on GitHub at https://github.com/meituan-longcat/LongCat-2.0 and on HF (MIT).

Review and scope

  • The blog and model card share a source (the same team, released the same day), so their technical details mutually corroborate each other; they also agree with AlphaSignal's independent analysis (Review 04) on the three LSA strategies, N-gram, and the three MOPD expert groups (AlphaSignal adds: approximately 128B N-gram parameters versus the official 135B, a minor difference in scope).

  • The blog does not name the chip vendor ("domestic compute chips"); the community infers Huawei Ascend 910C from the official acknowledgment of HCCL (Huawei's communication library) (see the r/LocalLLaMA discussion in Review 14), but this is an inference rather than official confirmation.

  • Scope: the blog is an engineering note rather than an independent evaluation. Official throughput/reliability figures (such as a 70%+ reduction in failure rate) have no third-party verification; the deployment approach targets very large clusters and offers individual developers mainly the SGLang cookbook as a reference.

  • Related documents: Review 01 (official benchmark table), Review 04 (AlphaSignal verification).

What this supports

  • The blog and model card share a source (the same team, released the same day), so their technical details mutually corroborate each other; they also agree with AlphaSignal's independent analysis (Review 04) on the three LSA strategies, N-gram, and the three MOPD expert groups (AlphaSignal adds: approximately 128B N-gram parameters versus the official 135B, a minor difference in scope).
  • The blog does not name the chip vendor ("domestic compute chips"); the community infers Huawei Ascend 910C from the official acknowledgment of HCCL (Huawei's communication library) (see the r/LocalLLaMA discussion in Review 14), but this is an inference rather than official confirmation.

What this does not support

  • Scope: the blog is an engineering note rather than an independent evaluation. Official throughput/reliability figures (such as a 70%+ reduction in failure rate) have no third-party verification; the deployment approach targets very large clusters and offers individual developers mainly the SGLang cookbook as a reference.
  • Related documents: Review 01 (official benchmark table), Review 04 (AlphaSignal verification).

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

LongCat official blog (longcat.chat) · Meituan LongCat team (official) · Original publication date 2026-06-30 · Site edit date 2026-09-20

Open original source

LongCat 2.0

Compare LongCat 2.0 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

LongCat 2.0: what changed, where to use it, and what the price misses

LongCat 2.0 combines 1M context, open weights, and low provider pricing with real questions about tooling, data terms, and operational cost.

Related reviews

LongCat-2.0 Official Model Card: Specifications and Official Benchmarks (Including Comparison Tables with Gemini/GPT-5.5/Claude Opus)The official model card is the primary authoritative source for judging LongCat-2.0's suitable tasks: it scores 59.5 on SWE-bench Pro, ahead of GPT-5.5 (58.6) and Gemini 3.1 Pro (54.2), and reaches 70.8 on Terminal-Bench 2.1. However, it trails GPT-5.5 and Claude Opus 4.8 on several benchmarks including BrowseComp, GPQA, and IFEval—in short, it is strong at coding and agent tasks, but not a leader in retrieval and general reasoning.eesel Independent Review: LongCat-2.0's Agent Reliability and Hard Blockers to Production DeploymentThis independent review separates LongCat-2.0 into two questions: "can the model complete Agent work?" and "can the product enter enterprise production?" Public user reports support it as an inexpensive, stable coding executor, but its context specifications, tool contract, and data-governance documentation are insufficient to pass a sensitive-data production review.OpenRouter Channel Data: LongCat-2.0 Pricing, Measured Performance, and Third-Party Benchmarks (Artificial Analysis)The OpenRouter page provides a third-party view beyond the official figures: LongCat-2.0 is listed at $0.30/$1.20 per 1M tokens (with a 60% discount at collection time), while the actual weighted transaction price for input was only $0.03872/M (88.9% cache-hit rate); throughput was P50 29 tok/s, three-day availability 99.93%, and tool-call error rate 0.90%, with real traffic mainly coming from Hermes Agent (7.77B tokens) and Claude Code (3.31B tokens).AI Profit Boardroom field test: LongCat 2.0 game-building test and same-task comparison with GLM 5.2The author's test reached a conclusion opposite to most community sentiment: LongCat 2.0's games were "playable but rough and buggy" (one build even showed a completely black screen), while GLM 5.2's outputs on the same tasks were "cleaner, smoother, and more polished." His recommendation was "worth playing with, not worth switching to" — a negative independent sample that should be read alongside positive evidence about LongCat 2.0.LongCat-2.0 API Platform Quick Start (Official Quick Start + Chat Completions Reference + Pricing)The LongCat Claude Code guide configures a compatible endpoint and keeps the first task in a disposable worktree.LongCat-2.0 Chat Template and Tool-Calling Configuration (Official Hugging Face Model Card)The official model card’s chat template and tool-call examples are converted into a local inference configuration check.Claude Code Integration with LongCat-2.0 (Official Documentation)The official LongCat integration guide configures a named client and keeps the first run observable and reversible.Hermes Agent Integration with LongCat-2.0 (Official Documentation + Nous Portal Free Entry)The official LongCat guide “Hermes Agent Integration with LongCat-2.0 (Official Documentation + Nous Portal Free Entry)” configures a named client and keeps the first run observable and reversible.