Adit_Yah says the same prompt produced completely different code results three days apart, with duration increasing from 80 minutes to 330 minutes and more than 4,500 lines of code; this is better treated as a signal for reproducing routing, version, or sampling variance than as evidence of continual learning.
Model: Ox Alpha; the client, provider, version snapshot, parameters, and complete prompt were not disclosed.
Comparison window: The author says the same task was run three days earlier and again that day; complete timestamps, commits, and model IDs for both runs were not published.
Task: The post's video shows a code/racing-scene task of the same type, but no repository, input files, or acceptance script were provided.
Reported outcomes: The current run took 330 minutes and produced more than 4,500 lines of code; the earlier run took 80 minutes. A commenter said the first video might be played at 1.5x speed; the author clarified that video speed and the track were separate issues.
| Metric | Visible post/reply content | Interpretation boundary |
|---|---|---|
| Prompt | The author says both runs used exactly the same prompt | The prompt was not published, so byte-level identity cannot be checked |
| Code volume | More than 4,500 lines in the new run | No diff, effective-code ratio, or indication of generated files |
| Duration | 330 minutes vs. 80 minutes | No wall-clock log or separation of waiting from human actions |
| Explanations discussed | Routing to a weaker model, service congestion, version change, and overthinking | The alternatives were not separated by a controlled experiment |
Save the complete prompt, repository commit, model ID, provider, client version, temperature/effort, tool permissions, and timestamp.
Repeat at least 5 runs on the same day to establish a sampling/service-variance baseline, then repeat the same group across multiple dates.
Save the complete tool trace, output tokens, generated files, errors, retries, total wall-clock time, and human intervention.
Fix or record video playback speed, hardware, track/input resources, and the acceptance script.
Compare final test pass rate, effective diff, rework time, and cost rather than only line count or video impression.
The original post supports the reproducibility signal that Ox Alpha may show substantial cross-date result differences during a short preview; it does not support “the model is continually learning” or “the model was definitely stronger that day.” Users who need stable results should record the model, provider, and time window in the experiment log.
No complete prompt, code, tests, model version, or request logs are provided.
The two runs had no same-day repetitions or randomized controls.
Video speed and task-scene details are disputed, so video duration cannot directly measure performance.
Ox Alpha