The author placed Gemini 3.8 Flash, GPT 5.6 Sol, Claude Opus 5, and Kimi K3 under the same brief and had each generate a 3D rocket-launch scene once. The author concluded that Flash did produce a runnable scene, but lost this comparison on visual geometry, smoke effects, and frame stability; the post also records that spending $2.30 and running in Goal Mode for 30 minutes did not save the result.
Tasks worth observing: Fast 3D/visual generation from the same brief, and creative tasks that require checking whether a scene runs and assessing the quality of the final output.
Tasks that should not be inferred directly: General 3D capability, all video-generation tasks, coding ability, or an overall model ranking; the original post contains only one task comparison.
Applicable model version: The original post explicitly says Gemini 3.8 Flash; no public snapshot, reasoning tier, or provider is given, so the result should not be extended to other Gemini versions or interfaces.
Applicable client, agent, or API: Not disclosed. The post mentions only Goal Mode and does not specify the product, toolchain, model invocation method, or permissions.
Recommended reasoning tier and parameters: Not disclosed; do not treat the 30 minutes mentioned in the post as a reusable default budget.
The author states that all four models used the same brief and that the task was to generate four 3D rocket launches; the full brief text was not made public.
The comparison models were Gemini 3.8 Flash, GPT 5.6 Sol, Claude Opus 5, and Kimi K3. The original post did not disclose each model's complete input, system prompt, tools, output specifications, or number of runs.
The direct observation for Flash was that the scene was "runnable," but its geometry looked toy-like; hard white cones stood in for smoke; and one frame became blank.
The author reports that this attempt consumed $2.30 and ran for 30 minutes in Goal Mode; the original post does not explain the fee basis, billing party, token usage, number of retries, or whether these two figures applied only to Flash.
"Lost" is the author's overall visual judgment of the four results, not a public scoring table or independently reviewed result. The replies are not included in this entry.
| Observation dimension | Public content in the original post | Limited judgment supported |
|---|---|---|
| Task | Same brief; four 3D rocket launches | At least one cross-model, same-task comparison was conducted by the user |
| Comparison models | Gemini 3.8 Flash, GPT 5.6 Sol, Claude Opus 5, Kimi K3 | The author's relative comparison set can be described, but strict A/B conditions cannot be reconstructed |
| Whether Flash completed the task | The author says it did not "fail" on 3D and produced a runnable scene | Flash can produce a runnable result; "runnable" does not mean the visual quality met the bar |
| Visual problems | Toy-like geometry; hard white cones replacing smoke; one blank frame | This output had clear problems with form, effects, and frame stability |
| Cost and duration | $2.30; Goal Mode for 30 minutes | The experience required a non-trivial investment, but the actual per-run bill or value for money cannot be calculated from it |
| Overall judgment | The author says Flash "lost" this four-model comparison | Supports only a subjective relative evaluation for this task, not a model ranking |
The clearest signal from this first-hand experience is the gap between "runnable" and "finished-output quality": Flash was not unable to generate a 3D scene, but this result showed low-detail geometry, an incorrect substitute for smoke, and a blank frame.
In the author's same-brief comparison, adding 30 minutes in Goal Mode and a $2.30 investment did not eliminate these problems either; for tasks that require stable frames and acceptable visual polish, simply extending runtime is not a sufficient fallback.
This is not a benchmark. The original post did not disclose the input, scoring criteria, run logs for each model, or result metrics, so "lost" can serve only as a retesting lead: focus on geometric detail, effects semantics, and cross-frame stability.
The evidence consists only of one X post by the author; there is no task set, repeated experiment, blinded evaluation process, success rate, or statistical confidence.
"The same brief" is the author's statement, but the complete brief, system prompts, and model invocation settings are not visible, so it is impossible to verify whether all four models were tested under exactly the same conditions.
The specific client, provider, model snapshot, context, sampling parameters, tool permissions, output resolution, video duration, and random seed were not disclosed.
$2.30 has no billing breakdown, and the 30 minutes are not broken down into waiting, inference, or manual-operation time; they cannot be extrapolated as the per-task cost or latency of Gemini 3.8 Flash.
The post includes a video of about 25 seconds, but this article does not treat the video display as an independent score; "toy-like geometry," "hard white cones," "one blank frame," and "lost" all come from the author's observations and judgments.
Opinions in the replies about "who won" were not included, to avoid presenting other users' subjective comments as the original author's experimental results.
Fix the same 3D rocket-launch brief, output duration, resolution, toolchain, model snapshot, and budget; record each model's complete input and system configuration.
Run Gemini 3.8 Flash, GPT 5.6 Sol, Claude Opus 5, and Kimi K3 multiple times in the same environment, separately recording successful generation, scene runnability, geometric detail, smoke/effects semantics, and blank frames.
Record Goal Mode duration, actual waiting time, number of retries, token/billing data, and manual correction time separately; do not use $2.30 or 30 minutes alone to represent cost and latency.
Use a blinded evaluation or a predefined scorecard containing at least "runnable," "complete across frames," "subject and effects match the brief," and "visual quality"; then report the mean, failed samples, and variance.
The original post explicitly lists the four models, the same brief, and four 3D rocket launches. The link is https://x.com/Its_lakshya_ai/status/2095525460629455002.
The visible observations for Gemini 3.8 Flash in the original post are that it generated a runnable scene, but the geometry looked toy-like, hard white cones replaced smoke, and one frame was blank.
The original post also states $2.30 and 30 min in Goal Mode, and says that this investment did not save the result; no billing, token, or run logs were disclosed.
The author's core judgment was: "Flash did not fail on 3D; it produced a runnable scene, it just lost." This should be understood as a personal judgment in a single-task visual comparison, not as a conclusion about general capability.
Gemini 3.8 Flash