The developer put DeepSeek V4 Flash 0731, DeepSeek V4 Pro, and LongCat 2.0 through the same "physics + coding" test (same task, same conditions). The result: "LongCat 2.0 surprised me the most" — it performed far more successfully than expected, and this was his first time encountering it on a free tier.
I put DeepSeek V4 Flash 0731, DeepSeek V4 Pro, and LongCat 2.0 through the same physics + coding test. All three models… This is a necessary excerpt; read the original source for full context.
This was a self-designed test by a single author. The task details, scoring criteria, and tier were not disclosed (only "same task, same conditions" was stated), and the video is evidence of the output. It supports a directional conclusion — "LongCat 2.0 exceeded expectations on a physics-and-coding hybrid task and can compete alongside the DS4 series" — but not a precise ranking.
The conclusion is directionally consistent with r/hermesagent's positioning (between DS4-Flash and DS4-Pro, review 06): LongCat 2.0 and the DeepSeek V4 series occupy the same capability band, with intense value competition (AlphaSignal notes that DS4-Pro is listed at a lower price, review 04).
The free-tier experience is consistent with the "OpenRouter/Nous Portal free endpoint" (prompt directory 03), showing that a free entry point is indeed available and usable in performance terms.
LongCat 2.0