One user used the same zero-shot creative prompt in Google Antigravity to have Gemini 3.1 Pro, 3.6 Flash, 3.7 Flash, and 3.8 Flash create a simple game combining Y City and Minecraft. The author felt that 3.7 and 3.8 Flash worked longer and more autonomously, but under this loose creative prompt, 3.6 Flash captured the intent more accurately; therefore, greater autonomy in 3.8 Flash does not necessarily mean a higher hit rate on creative intent.
Tasks suitable for observation: Creative Agent tasks that require a model to turn a loose, open-ended request into a runnable prototype, and version comparisons using the same prompt.
Tasks not suitable for direct inference: General game-development ability, code quality, long-context capability, speed, or production-grade success rate; the original post did not disclose detailed artifacts, run logs, or a scoring table.
Model versions covered: The author explicitly compared Gemini 3.8 Flash with Gemini 3.7 Flash, 3.6 Flash, and 3.1 Pro; this cannot be extended to other Gemini versions, reasoning tiers, or providers.
Applicable client, Agent, or API: Google Antigravity; the API model ID, tool configuration, permissions, context limits, and sampling parameters were not disclosed.
Recommended reasoning tier and parameters: Not provided; “working longer” describes the author's observation of session behavior and cannot be treated as a reusable time budget or effort configuration.
Environment: Google Antigravity; the author described this as an empirical test and compared the four models available at the time.
Input: The same zero-shot prompt asking the models to create a simple game combining Y City and Minecraft. The full prompt was not disclosed, and the author also said that detailed results could not be shared.
Comparison: Gemini 3.1 Pro, Gemini 3.6 Flash, Gemini 3.7 Flash, and Gemini 3.8 Flash were given the same task description; the original post did not provide the number of runs for each model or complete outputs.
Observation: The author believed that Gemini 3.7 and 3.8 Flash “remain working longer and autonomously”; however, for understanding and accuracy on the loose creative request, the author preferred Gemini 3.6 Flash and considered 3.1 Pro disappointing.
Supporting evidence: The only comment argued that stronger autonomy can sometimes cause a model to deviate from what the user actually asked for; this is the commenter's view and is not counted as a result from the original post.
| Observation dimension | What the original post publicly says | Limited judgment supported |
|---|---|---|
| Client | Google Antigravity | The conclusion applies only to a session in this client and cannot be directly equated with API behavior |
| Task | Create a simple Y City + Minecraft game with the same zero-shot prompt | A cross-version creative comparison exists for one task |
| Gemini 3.7 Flash | Like 3.8 Flash, the author felt that it worked longer and more autonomously | Supports only a directional observation about session behavior, not differences in duration or completion rate |
| Gemini 3.8 Flash | Worked longer and more autonomously, but the author did not consider it the most accurate at capturing the loose creative prompt | Autonomous execution and creative-intent hit rate are separate dimensions |
| Gemini 3.6 Flash | The author thought it had better understanding and accuracy on the loose creative prompt | Supports this author's subjective preference on this task, not a general ranking |
| Quantitative data | No token counts, elapsed times, scores, retries, complete outputs, or code diffs | Speed, cost, code quality, or success rate cannot be calculated |
This experience provides a signal different from existing code-repository reports: the task was not fixing a real repository, but turning a vague creative intent into a game prototype.
In the author's same-task comparison, Gemini 3.8 Flash and 3.7 Flash were stronger in working longer and more autonomously; however, the author awarded the ability to accurately understand the loose creative prompt to 3.6 Flash.
Therefore, for open-ended creative Agents, upgrade effects cannot be measured only by the number of runs, autonomous tool calls, or continued working time; goal intent, thematic integration, and whether the artifact matches user expectations should also be evaluated separately.
This was not a controlled evaluation: the original post provided no detailed prompt, independent result for each of the four models, scoring criteria, or logs. “3.6 understands the request better” can only serve as a hypothesis for retesting.
The author did not disclose detailed results for the game, providing only key takeaways; the text cannot establish whether the four versions' code was runnable, whether features were complete, or what the visual quality was.
The model snapshot, reasoning tier, system prompt, tool permissions, context length, temperature, run count, elapsed time, and token count were not disclosed. Therefore, “working longer” cannot be interpreted as higher server-side latency or a larger budget.
“The same zero-shot prompt” is the author's statement, but the prompt text and each model's context state are not visible, so identical conditions cannot be fully verified.
The 3.7 and 3.8 comparison is reported only as a combined description by the author, without per-model outputs. It cannot establish that 3.8 was better or worse than 3.7; it can only say that the author did not report an additional advantage for 3.8 in intent understanding.
The author provided the video link as a walkthrough, but this article does not derive additional conclusions from the video footage; a retest should preserve the video, code, and original session for each model.
The conclusion comes from one user, one task, and subjective judgment. It cannot be generalized to Gemini 3.8 Flash's overall creative ability or statistically significant differences from 3.6/3.7.
Fix the same Y City + Minecraft creative brief, system prompt, Antigravity version, model snapshot, tool permissions, and work-time limit.
Run Gemini 3.6 Flash, 3.7 Flash, and 3.8 Flash multiple times each; save the complete prompts, plans, tool calls, code, runnable builds, and final screenshots/videos.
Score “autonomy” and “creative hit” separately: record effective working time, tool calls, and premature stopping for the former; score thematic integration, core gameplay coverage, user-intent hit, runnability, and required manual rework time for the latter.
Predefine what counts as “deviating from the request,” such as the model replacing the theme, removing key gameplay, or expanding the scope without clarification; do not use output length as a substitute for intent hit.
When retesting 3.7 and 3.8, report results and failure samples for each run; if only the author's experiential conclusion is retained, label it as community-opinion rather than a benchmark score.
The original post explicitly says that the author used the same zero-shot prompt in Google Antigravity to compare Gemini 3.1 Pro, 3.6 Flash, 3.7 Flash, and 3.8 Flash: https://www.reddit.com/r/GeminiAI/comments/1wa2xwf/i_raised_all_four_models_currently_present_in/.
The original post limits the task to creating a simple game combining Y City and Minecraft, and says that detailed results could not be shared, providing only key takeaways.
The author's key observations were that 3.7 and 3.8 Flash worked longer and more autonomously; 3.6 Flash had better understanding and accuracy on the loose creative prompt; and 3.1 Pro was disappointing.
The author described the task as using “the exact same zero-shot prompt” and explicitly said that 3.6 Flash was more accurate with “a loose, creative prompt.” Both statements represent this individual comparison only and cannot replace an independent retest on a standardized task set.
Gemini 3.8 Flash