As a trillion-parameter model, V4 Pro meets Seed 2.0 Pro and Kimi K2.6 at the top of the first tier, taking first place among Chinese-developed models by a significant margin; its max tier, however, has substantially higher reasoning overhead.
V4 Flash, with a scale of 200B+ parameters, meets the similarly sized Hy3 at the bottom of the first tier. V4 Flash is cheaper, but the two have comparable total costs and trade wins and losses.
Instruction following: V4 Pro follows instructions reliably, performing consistently across multiple passes in complex contexts with multiple conditions. Its capability is very close to GPT-5.4 and significantly higher than Kimi K2.6. The max tier ignores instructions such as "do not overthink," while the high tier responds to such requests. V4 Flash's instruction-following capability is broadly on par with V4 Pro and slightly better than Hy3.
Complex reasoning: On multi-step, long-chain reasoning, V4 Pro matches K2.6 at the upper bound and is slightly weaker than GPT-5.4, but its performance is not stable enough—the max tier tends to overthink and randomly get stuck in local solutions. V4 Flash is affected in the same way; on moderately difficult tasks, it is actually less stable than Hy3.
Context hallucinations: V4 Pro does hallucinate, but not at a high level. Minor errors usually appear when the context contains a large amount of similar text. On long-text information-extraction tasks, information correctly extracted in the chain of thought can become distorted during later processing. V4 Flash's hallucination level is on par with V4 Pro, while both its upper bound and stability are better than Hy3's.
Pattern insight: V4 Pro does not demonstrate the level of insight expected from its parameter scale. In mathematical-symbol derivations and letter-pattern exploration, GPT-5.4 and Opus show genuine insight without requiring long reasoning, whereas V4 Pro relies heavily on inefficient exhaustive enumeration without pruning. Its reasoning length often approaches the prescribed limit, with a 50% probability of running long. V4 Flash inherits some degree of this insight but behaves randomly; overall, it trades wins and losses with Hy3.
Inefficient reasoning: The max tier is significantly less efficient at reasoning. At the same accuracy, the high tier usually consumes only one-half or even one-third as many tokens as max. Although max has higher accuracy on the hardest problems, the improvement does not match its token consumption. Compared with GPT-5.4's xhigh tier, GPT-5.4 can keep intelligence and token consumption close to a linear relationship.
Models released by DeepSeek are often SOTA in certain areas at the time of release, and their cost-effectiveness can remain a benchmark over the long term. As a committed advocate of open source, its technology-for-all approach means that even when its models are closed-source and paid, there is no shortage of users willing to pay.
DeepSeek V4 Flash