The original poster says the model was very fast when paired with DSH to develop a game engine, but completed far less work than Opus and forgot Markdown rules it had already read; this is a personal risk signal for complex Agent tasks and cannot be attributed solely to the model or DSH.
Tasks this can help assess: Rule following and task retention in complex code projects.
Tasks this should not be extrapolated to: General programming ability, success rate, cost, or a fair comparison with Opus.
Applicable model version: DeepSeek V4.1 Flash.
Test environment or client: DSH; the provider, specific configuration, and project scale were not specified.
Reasoning tier and parameters: Not specified.
There was no task set, full prompt, parameter specification, number of repetitions, or acceptance criteria. The original poster says they spent $10 on game-engine development, but provides no error list, completion rate, or subsequent OpenCode results, and suspects DSH may also have had an effect. The original poster emphasized that the key question was how much work was actually completed, rather than the number of tokens; this is only their evaluation criterion, with no itemized record or review results.
The original poster says the model thinks and generates tokens quickly, but quickly forgets context during actual work; even when the rules are in a Markdown file it has read, it still does things that are explicitly forbidden. A Hermes Agent user self-reported that randomly selected Obsidian note rules loaded at the start of a session were still remembered after about 600k tokens. The tasks and configurations differed, so the latter is not a counterexample.
| Item | Information visible in the original post |
|---|---|
| Model | DeepSeek V4.1 Flash |
| Client | DSH |
| Task | Building a game engine |
| Cost | The original poster says they spent $10 |
| Observed behavior | Fast; forgetting rules; deviating from the task and performing forbidden actions |
| Parameters and sample | Not specified; one person's experience |
The post suggests that, in long-running development, speed cannot replace rule retention and task convergence; however, it cannot determine whether the problem came from the model, DSH, prompt design, or project complexity. Controlled reproduction and human acceptance are still required. The Hermes claim about 600k tokens is another person's self-report. Subscription token counts in the comments are not treated as factual evidence.
Fix DSH, the model version, the game-engine task, and the Markdown rules; record context length, the number of violations, effective changes, and acceptance results. After repeating the test, compare OpenCode or Hermes as separate conditions.
DeepSeek V4.1 Flash