A user tested a $2 package containing 50 million tokens: a single task used 27 million prompt tokens, but only 570,000 tokens were deducted because of cache hits. The official rule — "cache hits are not billed; only misses and output count" — expands effective usage to roughly 2 billion tokens in a real agent workflow, making this the most counterintuitive part of LongCat's pricing.
Motivation: The user was intrigued by Meituan's LongCat 2.0 but hesitant because there were not many benchmarks. Everyone was praising Owl Alpha, though, and "it turned out to be the same model."
Purchase: The official 50M-token package costs $2 (the actual charge was €0.86, "not sure why"); installing AliPay on a phone was required and took about 15 minutes.
Key billing rule: The 50M tokens apply only to cache misses or output; cache hits are free!
Test: One task used 27 million tokens (apparently the total prompt/context volume), while the bill deducted only 570,000 tokens.
Conclusion: "At this rate I guess I can use around 2 billion effective tokens lol."
This billing model matches the official pricing page (the pricing entry in prompt directory 01: Cached Input ¥0.10 → discounted ¥0.04, far below Uncached ¥5 → ¥2). It is also consistent with OpenRouter's 88.92% cache-hit rate and actual weighted input price of $0.03872/M (review 03), confirming that caching is the core lever behind LongCat's value.
Best-fit tasks: Agents that repeatedly reread the same context (coding loops and long-document research) benefit most (the same conclusion as AlphaSignal); one-shot prompts receive almost none of the caching benefit.
Limits: These prices were promotional ($2 package; the $4.9 package had a 67% discount), and the current price should be checked on the platform. AliPay registration is a barrier for users in China (other commenters mentioned needing a phone number/AliPay; international users can bypass it through OpenRouter/Nous Portal channels; see prompt directory 03/06).
LongCat 2.0