Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

LongCat 2.0 · Community source · Personal experience

AlphaSignal Deep Dive: Owl Alpha's True Identity and a Reality Check on LongCat-2.0's Official Claims

This deep-dive review, published the day after launch, establishes the most important background fact — the anonymous free model "Owl Alpha," which ran on OpenRouter for two months, was LongCat-2.0 (processing approximately 10 trillion tokens per month and ranking first on the Hermes Agent leaderboard) — while breaking down the official claims one by one: the narrow SWE-bench Pro win over GPT-5.5 was an official self-test, the weights were still only "planned" when the review was published, and the open-source fiel

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model/version
LongCat-2.0; source date: 2026-07-01.
Harness/task
This deep-dive review, published the day after launch, establishes the most important background fact — the anonymous free model "Owl Alpha," which ran on OpenRouter for two months, was LongCat-2.0 (processing approximately 10 trillion tokens per month and ranking first on the Hermes Agent leaderboard) — while breaking down the official claims one by one: the narrow SWE-bench Pro win over GPT-5.5 was an official self-test, the weights were still only "planned" when the review was published, and the open-source fiel
Sample/gaps
Limitations noted: The article is analytical commentary rather than a controlled evaluation. Its figures are relayed from official and platform sources, but key data such as "10 trillion tokens/month" and "GLM-5.2 62.1" have no independent original-source links (the author says source links are in the first reply); treat them as second-hand when citing.; It is best read alongside the official documents (Reviews 01 and 02): the official documents provide the full details, while this article provides a critical framework.

Key data and applicable tasks

One-sentence takeaway

This deep-dive review, published the day after launch, establishes the most important background fact — the anonymous free model "Owl Alpha," which ran on OpenRouter for two months, was LongCat-2.0 (processing approximately 10 trillion tokens per month and ranking first on the Hermes Agent leaderboard) — while breaking down the official claims one by one: the narrow SWE-bench Pro win over GPT-5.5 was an official self-test, the weights were still only "planned" when the review was published, and the open-source field's true rival was GLM-5.2 (62.1 vs. 59.5).

Core facts and arguments (key points from the original)

Owl Alpha background

  • Owl Alpha ran anonymously on OpenRouter for two months; when developers called it, they did not know the vendor or internal structure. At launch it was processing approximately 10 trillion tokens per month, more than any model on OpenRouter's Hermes Agent leaderboard.

  • The only evidence not self-reported by the official source was the two months of anonymous usage: "This is the strongest piece of evidence in the entire launch, and the only thing Meituan did not generate itself."

  • Risk disclosure: Owl Alpha's OpenRouter listing disclosed that prompts and completions might be logged by the provider to improve the model — free users unknowingly supplied training signals for two months.

Technical analysis (cross-checked against the official blog)

  • Zero-Computation Experts: simple tokens (variable names, parentheses) take an almost empty computation path, while difficult tokens go through the full expert stack; routing is per token rather than per request.

  • LSA: the official source acknowledged that the indexer itself became a new bottleneck (scoring every layer and token remains quadratic, and the memory-access pattern does not match hardware prefetching); SI/CLI/HI are the three corresponding fixes. The risk is that "when incorrect routing sends a token that needs full reasoning down a cheap path, it will fail silently," and no public figure explains how often this happens.

  • N-gram Embedding: n-gram size 5, approximately 128B parameters (official figure: 135B), and an effective vocabulary expansion of approximately 900x; MoE + N-gram produces effective sparsity of approximately 97%, with N-gram accounting for <10% (returns diminish after >30%).

  • MOPD: post-training is split into three expert groups — Agent / Reasoning / Interaction — with a gating network routing by task at inference time — "probably the most genuinely novel part"; most labs only merge weights by averaging.

Hardware and training

  • The official source says this is the first trillion-parameter model trained entirely on a domestic cluster (approximately 50,000 accelerators, no Nvidia), but does not name the chip vendor; the official acknowledgment of Huawei's HCCL communication library and independent estimates point to Huawei Ascend 910C.

  • Official engineering claims: monthly hardware failure rate reduced by 70%+ and 35T tokens processed with no rollback or unrecoverable loss spikes — "achieving this on immature non-CUDA hardware is a real accomplishment."

Benchmark verification

  • In the official self-test table, the only model LongCat-2.0 consistently beats is Gemini 3.1 Pro; against GPT-5.5 it only narrowly wins SWE-bench Pro (59.5 vs. 58.6) and trails across the rest (FORTE 73.2 vs. 77.8, RWSearch 78.8 vs. 85.3, BrowseComp 79.9 vs. 84.4).

  • Claude Opus 4.8 is a bigger threat to LongCat-2.0: on the same SWE-bench Pro test it scores 69.2 (nearly 10 points ahead), while Opus 4.7 also scores 64.3.

  • Three claims thin out under line-by-line review:

    1. The weights had not been released at the time — GitHub explicitly said "Model weights coming soon, stay tuned" (only 76 stars/4 forks 10 hours after launch); media reports that "Meituan open-sourced" it described a plan rather than a release. (Note: as of 2026-08-18, the weights have been released on HF, so this point is outdated.)

    2. The strongest open-source rival, GLM-5.2, was absent from the official comparison chart: GLM-5.2 scored 62.1 > 59.5 on the same SWE-bench Pro test, and its MIT-licensed weights were downloadable two weeks earlier.

    3. The official chart used the company's own harness and scoring throughout, and noted that "problematic tasks [were] corrected"; at publication, there were no independent third-party scores from Artificial Analysis, Scale, or anywhere else.

Final conclusion (original text)

  • Cache hits are most cost-effective for "Agents that repeatedly reread the same context" (single-repository coding Agents, single-long-document research Agents); one-shot prompts get almost no benefit.

  • DeepSeek V4-Pro's list price is still lower than LongCat-2.0's; use GLM-5.2 to verify the benchmarks (downloadable weights, higher SWE-bench Pro score); browsing, multi-step tool chains, and in-task judgment remain areas where closed frontier models lead.

Review and scope

  • The article was published on 2026-07-01 (the day after launch). "Weights not released" and some pricing data ($0.69/$2.78 per 1M) are outdated: the weights are now available on HF (Prompt Directory 02), official current pricing is ¥5/¥20 (Prompt Directory 01), and OpenRouter's current price is $0.30/$1.20 (Review 03).

  • The article is analytical commentary rather than a controlled evaluation. Its figures are relayed from official and platform sources, but key data such as "10 trillion tokens/month" and "GLM-5.2 62.1" have no independent original-source links (the author says source links are in the first reply); treat them as second-hand when citing.

  • It is best read alongside the official documents (Reviews 01 and 02): the official documents provide the full details, while this article provides a critical framework.

What this supports

  • The article was published on 2026-07-01 (the day after launch). "Weights not released" and some pricing data ($0.69/$2.78 per 1M) are outdated: the weights are now available on HF (Prompt Directory 02), official current pricing is ¥5/¥20 (Prompt Directory 01), and OpenRouter's current price is $0.30/$1.20 (Review 03).
  • The article is analytical commentary rather than a controlled evaluation. Its figures are relayed from official and platform sources, but key data such as "10 trillion tokens/month" and "GLM-5.2 62.1" have no independent original-source links (the author says source links are in the first reply); treat them as second-hand when citing.

What this does not support

  • The article is analytical commentary rather than a controlled evaluation. Its figures are relayed from official and platform sources, but key data such as "10 trillion tokens/month" and "GLM-5.2 62.1" have no independent original-source links (the author says source links are in the first reply); treat them as second-hand when citing.
  • It is best read alongside the official documents (Reviews 01 and 02): the official documents provide the full details, while this article provides a critical framework.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X (Twitter) Articles (AlphaSignal) · AlphaSignal (@AlphaSignalAI, AI industry news, 300,000+ subscribers) · Original publication date 2026-07-01 · Site edit date 2026-09-20

Open original source

LongCat 2.0

Compare LongCat 2.0 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

LongCat 2.0: what changed, where to use it, and what the price misses

LongCat 2.0 combines 1M context, open weights, and low provider pricing with real questions about tooling, data terms, and operational cost.

Related reviews

OpenRouter Channel Data: LongCat-2.0 Pricing, Measured Performance, and Third-Party Benchmarks (Artificial Analysis)The OpenRouter page provides a third-party view beyond the official figures: LongCat-2.0 is listed at $0.30/$1.20 per 1M tokens (with a 60% discount at collection time), while the actual weighted transaction price for input was only $0.03872/M (88.9% cache-hit rate); throughput was P50 29 tok/s, three-day availability 99.93%, and tool-call error rate 0.90%, with real traffic mainly coming from Hermes Agent (7.77B tokens) and Claude Code (3.31B tokens).LongCat-2.0 Official Model Card: Specifications and Official Benchmarks (Including Comparison Tables with Gemini/GPT-5.5/Claude Opus)The official model card is the primary authoritative source for judging LongCat-2.0's suitable tasks: it scores 59.5 on SWE-bench Pro, ahead of GPT-5.5 (58.6) and Gemini 3.1 Pro (54.2), and reaches 70.8 on Terminal-Bench 2.1. However, it trails GPT-5.5 and Claude Opus 4.8 on several benchmarks including BrowseComp, GPQA, and IFEval—in short, it is strong at coding and agent tasks, but not a leader in retrieval and general reasoning.LongCat-2.0 Official Technical Blog: Architecture, Training on Domestic Compute, and Inference Deployment (Release Notes)The official technical blog provides the complete technical foundation for LongCat-2.0 (LSA sparse attention, N-gram Embedding, 6D parallel training on domestic compute, and prefill-decode disaggregated deployment), making it useful for assessing the model's intended long-context and Agent capabilities, as well as reproducing the official benchmarks and deployment path.eesel Independent Review: LongCat-2.0's Agent Reliability and Hard Blockers to Production DeploymentThis independent review separates LongCat-2.0 into two questions: "can the model complete Agent work?" and "can the product enter enterprise production?" Public user reports support it as an inexpensive, stable coding executor, but its context specifications, tool contract, and data-governance documentation are insufficient to pass a sensitive-data production review.LongCat-2.0 API Platform Quick Start (Official Quick Start + Chat Completions Reference + Pricing)The LongCat Claude Code guide configures a compatible endpoint and keeps the first task in a disposable worktree.LongCat-2.0 Chat Template and Tool-Calling Configuration (Official Hugging Face Model Card)The official model card’s chat template and tool-call examples are converted into a local inference configuration check.Claude Code Integration with LongCat-2.0 (Official Documentation)The official LongCat integration guide configures a named client and keeps the first run observable and reversible.Official Account Showcase: Five “One-Prompt Generation” Creative Projects (Voxel/3D/CG/Landing Page/Mini-game)Source “Official Account Showcase: Five “One-Prompt Generation” Creative Projects (Voxel/3D/CG/Landing Page/Mini-game)” is organized as an executable task guide; its environment, inputs, and acceptance boundary follow the source.