Model ID: deepseek/deepseek-v4-flash (fast hybrid-attention reasoning)
Command: cmd --model deepseek/deepseek-v4-flash
Intelligence index: 52
Output speed: 115.9 tok/s
Pricing: input $0.22 /M; output $0.66 /M; cache reads $0.007 /M; Agent-loop cost $0.07 /M (in)
Context: 1M tokens; release: 2026-07-31
| Model | Intelligence | Coding | Speed | Input $/M | Output $/M | Blended $/M | Context |
|---|---|---|---|---|---|---|---|
| Muse Spark 1.2 Contributor | 56.8◆ | 72.2◆ | —◆ | $0.10◆ | $0.20◆ | $0.13◆ | 1.05M◆ |
| GPT-5.6 Luna | 52.3◆ | 71.4◆ | 159.5◆ | $0.20◆ | $1.20◆ | $0.45◆ | 1.05M◆ |
| Gemini 3.5 Flash | 52 | 70.1 | — | $1.50 | $9 | $3.38 | 1M |
| DeepSeek V4 Flash (latest) | 52 | 69.1 | 115.9 | $0.22 | $0.66 | $0.33 | 1M |
(◆ = best value in each column for that row)
Model: deepseek-v4-flash-text-to-text-0731 (also available: the web-search variant deepseek-v4-flash-web-search-0731)
Positioning: fast, affordable text generation, summarization, writing, and automation tasks, suitable for high-frequency calls, batch content processing, intelligent assistants, and scalable LLM workflows
OpenAI-compatible API: https://api.flaq.ai/api/v1/chat/completions, with streaming output support (text/event-stream)
Provides API pricing and benchmark information (specific figures are subject to the page's real-time values); X user @LeeLeepenkman says that deepseek-v4-flash-latest through the OpenRouter channel is still a “crazy deal” (an unusually good bargain)
Third-party platforms broadly confirm: a 1M-token context window, a 2026-07-31 release, an intelligence index of approximately 50–52, and an output speed of 100+ tok/s
Command Code's blended pricing (Blended $0.33/M) differs from both BenchLM's pricing and the official pricing, reflecting differences in each platform's assumptions about cache-hit rates
Pricing and benchmark scores may change as model versions iterate; when citing them, it is advisable to use the official DeepSeek API documentation and the platforms' real-time pages
DeepSeek V4 Flash