Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityGemini 3.8 Flash

Gemini 3.8 Flash: Four-Model 3D Rocket Launch Comparison and Generation Boundaries

Original source

X.com (Twitter)

AuthorLakshya (@Itslakshyaai)

Source date2026-09-03

Tabbit curation2026-09-08

Read original

One-sentence takeaway

The author placed Gemini 3.8 Flash, GPT 5.6 Sol, Claude Opus 5, and Kimi K3 under the same brief and had each generate a 3D rocket-launch scene once. The author concluded that Flash did produce a runnable scene, but lost this comparison on visual geometry, smoke effects, and frame stability; the post also records that spending $2.30 and running in Goal Mode for 30 minutes did not save the result.

Use cases

  • Tasks worth observing: Fast 3D/visual generation from the same brief, and creative tasks that require checking whether a scene runs and assessing the quality of the final output.

  • Tasks that should not be inferred directly: General 3D capability, all video-generation tasks, coding ability, or an overall model ranking; the original post contains only one task comparison.

  • Applicable model version: The original post explicitly says Gemini 3.8 Flash; no public snapshot, reasoning tier, or provider is given, so the result should not be extended to other Gemini versions or interfaces.

  • Applicable client, agent, or API: Not disclosed. The post mentions only Goal Mode and does not specify the product, toolchain, model invocation method, or permissions.

  • Recommended reasoning tier and parameters: Not disclosed; do not treat the 30 minutes mentioned in the post as a reusable default budget.

Test environment, inputs, and observations

  • The author states that all four models used the same brief and that the task was to generate four 3D rocket launches; the full brief text was not made public.

  • The comparison models were Gemini 3.8 Flash, GPT 5.6 Sol, Claude Opus 5, and Kimi K3. The original post did not disclose each model's complete input, system prompt, tools, output specifications, or number of runs.

  • The direct observation for Flash was that the scene was "runnable," but its geometry looked toy-like; hard white cones stood in for smoke; and one frame became blank.

  • The author reports that this attempt consumed $2.30 and ran for 30 minutes in Goal Mode; the original post does not explain the fee basis, billing party, token usage, number of retries, or whether these two figures applied only to Flash.

  • "Lost" is the author's overall visual judgment of the four results, not a public scoring table or independently reviewed result. The replies are not included in this entry.

Raw data

Observation dimensionPublic content in the original postLimited judgment supported
TaskSame brief; four 3D rocket launchesAt least one cross-model, same-task comparison was conducted by the user
Comparison modelsGemini 3.8 Flash, GPT 5.6 Sol, Claude Opus 5, Kimi K3The author's relative comparison set can be described, but strict A/B conditions cannot be reconstructed
Whether Flash completed the taskThe author says it did not "fail" on 3D and produced a runnable sceneFlash can produce a runnable result; "runnable" does not mean the visual quality met the bar
Visual problemsToy-like geometry; hard white cones replacing smoke; one blank frameThis output had clear problems with form, effects, and frame stability
Cost and duration$2.30; Goal Mode for 30 minutesThe experience required a non-trivial investment, but the actual per-run bill or value for money cannot be calculated from it
Overall judgmentThe author says Flash "lost" this four-model comparisonSupports only a subjective relative evaluation for this task, not a model ranking

Conclusion

  1. The clearest signal from this first-hand experience is the gap between "runnable" and "finished-output quality": Flash was not unable to generate a 3D scene, but this result showed low-detail geometry, an incorrect substitute for smoke, and a blank frame.

  2. In the author's same-brief comparison, adding 30 minutes in Goal Mode and a $2.30 investment did not eliminate these problems either; for tasks that require stable frames and acceptable visual polish, simply extending runtime is not a sufficient fallback.

  3. This is not a benchmark. The original post did not disclose the input, scoring criteria, run logs for each model, or result metrics, so "lost" can serve only as a retesting lead: focus on geometric detail, effects semantics, and cross-frame stability.

Scope and limitations

  • The evidence consists only of one X post by the author; there is no task set, repeated experiment, blinded evaluation process, success rate, or statistical confidence.

  • "The same brief" is the author's statement, but the complete brief, system prompts, and model invocation settings are not visible, so it is impossible to verify whether all four models were tested under exactly the same conditions.

  • The specific client, provider, model snapshot, context, sampling parameters, tool permissions, output resolution, video duration, and random seed were not disclosed.

  • $2.30 has no billing breakdown, and the 30 minutes are not broken down into waiting, inference, or manual-operation time; they cannot be extrapolated as the per-task cost or latency of Gemini 3.8 Flash.

  • The post includes a video of about 25 seconds, but this article does not treat the video display as an independent score; "toy-like geometry," "hard white cones," "one blank frame," and "lost" all come from the author's observations and judgments.

  • Opinions in the replies about "who won" were not included, to avoid presenting other users' subjective comments as the original author's experimental results.

Reproduction notes

  1. Fix the same 3D rocket-launch brief, output duration, resolution, toolchain, model snapshot, and budget; record each model's complete input and system configuration.

  2. Run Gemini 3.8 Flash, GPT 5.6 Sol, Claude Opus 5, and Kimi K3 multiple times in the same environment, separately recording successful generation, scene runnability, geometric detail, smoke/effects semantics, and blank frames.

  3. Record Goal Mode duration, actual waiting time, number of retries, token/billing data, and manual correction time separately; do not use $2.30 or 30 minutes alone to represent cost and latency.

  4. Use a blinded evaluation or a predefined scorecard containing at least "runnable," "complete across frames," "subject and effects match the brief," and "visual quality"; then report the mean, failed samples, and variance.

Original evidence and data

  • The original post explicitly lists the four models, the same brief, and four 3D rocket launches. The link is https://x.com/Its_lakshya_ai/status/2095525460629455002.

  • The visible observations for Gemini 3.8 Flash in the original post are that it generated a runnable scene, but the geometry looked toy-like, hard white cones replaced smoke, and one frame was blank.

  • The original post also states $2.30 and 30 min in Goal Mode, and says that this investment did not save the result; no billing, token, or run logs were disclosed.

Source excerpt or observation (for compliant short quotation only)

The author's core judgment was: "Flash did not fail on 3D; it produced a runnable scene, it just lost." This should be understood as a personal judgment in a single-task visual comparison, not as a conclusion about general capability.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Gemini 3.8 Flash

Use and compare models in Tabbit

Gemini 3.8 Flash

Related reviews

OfficialGoogle Blog (The Keyword)2026-09-02

Gemini 3.8 Flash: Google’s Official Benchmarks and Reproduction Boundaries

MediaArtificial Analysis (official model pages, methodology, and release article; the official X account was used to discover and cross-check the release post)2026-09-02

Gemini 3.8 Flash: Artificial Analysis Intelligence, Speed, Pricing, and Latency

MediaAI IQ (AIIQ, Liberated Software LLC)2026-09-02

Gemini 3.8 Flash: AI IQ Capability Benchmarks and Task Boundaries

MediaVals AI2026-09-05

Vals AI Finance Agent v2: Professional Finance Agent Benchmark for Gemini 3.8 Flash

Gemini 3.8 Flash

Related prompts

OfficialGoogle AI for Developers / Google DeepMind2026-09-02

Gemini 3.8 Flash: Google’s Official Model Parameters and API Configuration

OfficialGoogle AI for Developers2026-06-10

Gemini 3.8 Flash: Google's Official Structured Prompting and Agent Workflow

OfficialGoogle AI for Developers

Gemini 3.8 Flash: Google's Official Function-Calling Configuration and Tool Workflow

OfficialGoogle AI for Developers2026-09-02

Gemini 3.8 Flash: Google's Official Structured Output Configuration