LeMi compared the DeepSeek, OMP, and OpenCode frameworks using the same Ox Alpha, the same Interstellar Ranger RF-31D reference image, and the same prompt, suggesting that a harness may substantially affect results, but the post does not publish a quantitative score.
Model: The same Ox Alpha in all three runs; the post does not provide a model ID, provider, or version snapshot.
Input: The same Interstellar Ranger RF-31D reference image + exactly the same single prompt.
Comparison harnesses: DeepSeek, OMP, and OpenCode.
Task: Reproduce the object/scene from the reference image; the full task text was not published in the post.
Result artifact: A 12-second video; code, screenshots, console output, and acceptance criteria were not public.
Author preference: In a reply, the author said they liked the DeepSeek harness best, without explaining a scoring rubric.
| Item | Visible information | Boundary |
|---|---|---|
| Model control | Ox Alpha was run in all three harnesses | This does not prove that system prompts, tools, context, and parameters were identical |
| Input control | The same reference image and same single prompt | The prompt text was not published, so byte-level identity cannot be reproduced |
| Output | A result video was public | No source code, tests, rendering errors, or human score |
| Subjective preference | The author preferred the DeepSeek harness | Personal preference, not a model ranking |
Fix the same Ox Alpha model ID, provider, client version, system prompt, temperature/effort, tool schema, and timeout.
Copy the same reference image and save its SHA-256; save the complete prompt and initial context.
Clear the history in each of the three harnesses, changing only the harness, and record tool calls, file diffs, first token, total duration, and errors.
Define a uniform evaluation: semantic object, geometry/layout, runnable rendering, console errors, resource loading, and human rework time.
Repeat at least 3 times in randomized order, reporting model differences, harness differences, and service-load differences separately.
The main value of this case is separating “model capability” from “harness effects”: the same model output cannot be compared independently of client tools, context, and system prompt. The current evidence supports only a three-harness follow-up test; it does not support a quality ranking of DeepSeek, OMP, or OpenCode.
One run, no prompt text, no code, and no scoring script.
The three harnesses may have different hidden configurations.
A video can show an outcome but cannot establish reproducible code, no console errors, or long-term stability.
Ox Alpha