This discussion supports a narrow conclusion: marty4286 said V4.1 Flash had replaced Luna Max, cutting runs that had taken about 1 hour to 20–40 minutes; similar tasks with Sol high took about 15–30 minutes and cost less. But the task, inputs, client, parameters, and acceptance criteria were not disclosed, so this cannot be treated as a controlled comparison of speed or capability.
Tasks it can help assess: the iteration speed of coding agents working on complex code, budget routing, and workflows in which V4.1 Flash handles batch implementation while a stronger model reviews the work.
Tasks it should not be generalized to: general success rates for complex logic or multi-layer abstraction reasoning, capability rankings against Astra/Opus/Fable, stable costs, or production reliability.
Applicable model version: Both the original post and the comments identify the model as DeepSeek V4.1 Flash; the API ID, provider, and build version were not specified.
Test environment or client: The author mentioned Codex usage; the specific client, project, and harness were not specified.
Reasoning tier and parameters: Sol was set to high; the tiers for V4.1 Flash and Luna Max were not specified.
The original poster only asked whether anyone had compared these models in a real coding project. The comments are subjective accounts from multiple users, with no fixed task set, identical inputs, repeated trials, logs, or independent acceptance checks. marty4286 did not define what “equivalent tasks” meant, so the timings can only be treated as a personal workflow record.
marty4286 reported that a run taking Luna Max about 1 hour took V4.1 Flash 20–40 minutes; a similar task with Sol high took 15–30 minutes.
He considered V4.1 Flash “quite good” on its own, but prone to going off track and displaying worse bad habits than V4. Sol, Astra, and Fable could constrain it with instructions, but he was not yet willing to let it run fully unattended.
Some comments said it was far behind Opus or Sol on hard-core coding, while others considered it close enough for everyday tasks, showing that the task difficulty changes the conclusion. RealestReyn used a division of labor in which V4.1 Flash handled batch coding and Astra reviewed the work, and said the review had just found 5 bugs.
| Object | Information visible in the original post |
|---|---|
| V4.1 Flash | 20–40 minutes; replaced Luna Max; specific task not specified |
| Luna Max | About 1 hour for a similar run |
| Sol high | 15–30 minutes for a similar task; the author said it was cheaper |
| Discussion size | 1 post; 50 visible comments |
This post is better suited to supporting the field signal that “V4.1 Flash may improve coding iteration throughput and work well as a low-cost execution model.” It cannot prove that V4.1 Flash reaches Astra, Opus, or Fable on complex tasks. The comments’ reposts or paraphrases of the DeepSeek technical report’s point that “high-difficulty tasks still show a gap” are included only as discussion context and are not re-entered as an independent benchmark for this article; none of the timings is an experimental result from the same task with the same input.
Hold the same repository, prompts, tools, reasoning tier, and acceptance criteria constant. Record wall-clock time, number of turns, token/quota usage, cost, effective changes, and bugs found in review for each run. Repeat multiple times before comparing V4.1 Flash, Luna Max, and Sol high.
DeepSeek V4.1 Flash