Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Luna · Community source · Personal experience

Thoughts after using GPT-5.6 Luna for 48 hours

This evidence note covers “Thoughts after using GPT-5.6 Luna for 48 hours” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
Reddit, r/hermesagent; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.

Key data and applicable tasks

Summary

This 48-hour experience report contrasts with the positive reviews from the Codex community: using GPT-5.6 Luna high in Hermes agent for personal-assistant tasks, the author found it “smart but slow,” prone to repeated iteration, quick to consume quota, occasionally likely to miss explicit instructions, and more conservative in credential and CAPTCHA scenarios. The author's conclusion is that it is better suited to larger coding projects than as the default model for a lightweight personal assistant.

Key observations

  • In personal-assistant tasks, the author observed that Luna high often makes 5–6 iterations, while DeepSeek V4 Flash sometimes finishes in one.

  • The author says Luna consumes more than 20% of the Plus weekly quota each day, compared with about 9–11% for DeepSeek V4 Flash. These are personal-account and personal-workflow figures and should not be extrapolated to API pricing.

  • Direction-following is inconsistent: existing skills run normally, but the author says Luna occasionally ignores memory, soul directives, or part of a task.

  • It is more conservative with password and CAPTCHA operations; this is a security boundary and should not simply be treated as a lack of model capability.

  • Long-context compression appears frequently; the author saw prompts indicating that approximately 319k tokens were approaching the context/output limit.

Assessment

This article is a useful reminder not to choose a personal-assistant model based solely on coding benchmarks. Luna may offer a price advantage for high-frequency, coding, and verifiable tasks, but in personal-assistant use, tool calls, and long sessions, users should additionally measure completion rate, iteration count, quota consumption, and the rate of safety blocks.

Original article

The following is the body of the post extracted through the Tabbit international app; Reddit navigation, ads, and the community footer have been omitted.


Thoughts after using GPT 5.6 (Luna) for 48 hours

I've been using hermes agent for about 3 weeks now, and my usage has been mostly as a personal digital assistant: e-mails, package tracking, garden camera analysis, schedules and so on, with the occasional bash/php script creating and linux administration task.

For the past 2 weeks I've been testing different "low-cost" models, like deepseek-v4-flash (xhigh), minimax-m2.7 and many others.

My favorite model has been deepseek-v4-flash, but for the last 2 days I've been daily driving gpt-5.6-luna (high) using a Plus, and here are my findings:

  1. It's smart but slow: It knows how to solve problems, but most times it will go in circles, doing things in small steps instead of doing it fully once. Deepseek will do some things one-shot, while gpt-5.6-luna will think a lot and do 5–6 iterations.

  2. It's much more expensive: I've been using deepseek's model through OllamaCloud subscription, and a full day using hermes agent on deepseek-v4-flash with the occasional bigger model (for vision for example) will use 9-11% of my weekly allowance. With gpt-5.6-luna it's 20%+ per day, so my weekly allowance will be enough for 4–6 days only.

  3. It's not that good at following directions: All my skills would run seamlessly with deepseek or minimax, but with gpt-5.6-luna I had to go back to issues that were fixed days ago (and were explicit at the prompt). It ignores memory and soul directives as well, and sometimes will "forget" to do something, or even ignore part of the task completely.

  4. It's too "correct": It will not insert passwords in the browser (even after explicit authorization of the credentials' owner: me!) and you don't even think about asking for it to click a single captcha.

  5. 372k context is not enough?: Been seeing lots of "Pre-API compression: ~319,008 tokens near the context/output limit. Compacting before the next model call." lately.

The verdict is: at this time it's not a good model for me. It's pretty good for coding bigger projects for sure, but for light personal assistant use it's too expensive and dumb/stubborn.

Next week I'll subscribe to opencode's go $5 offer and test mimo-2.5 to let you guys know how it goes.

terra high for me has been the sweet spot. Sol its too much, luna too unpredictable.

It is slow and the password refusals and it silently inserting guardrails into things gets really really annoying. But it’s also like a dog after a bone give it a problem and it will keep going at it until it can fix it. That said I prefer ds4 it just does what I say do with out the push back.

What this supports

  • The 48-hour account calls Luna high smart but slow, often needing 5–6 iterations, missing instructions, and using over 20% of a Plus week.

What this does not support

  • It is one account’s experience; tasks, baselines, API records, and a version snapshot are missing, so the roughly 319k compaction cannot generalize to all accounts.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit, r/hermesagent · u/elzerouno · Original publication date Unknown · Site edit date 2026-09-20

Open original source

GPT-5.6 Luna

Compare GPT-5.6 Luna in Tabbit

Download the Tabbit client to check model access

Related reviews

Agents on Rails: 8 Models, 21 Atomic TasksThis evidence note covers “Agents on Rails: 8 Models, 21 Atomic Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Luna Is Really Underrated: Codex User ExperienceThis evidence note covers “GPT-5.6 Luna Is Really Underrated: Codex User Experience” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.5.6 Luna Extra High is the work horse I neededThis evidence note covers “5.6 Luna Extra High is the work horse I needed” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Luna Reddit Codex Quota and Cache Cost: A Hands-on MeasurementThis evidence note covers “GPT-5.6 Luna Reddit Codex Quota and Cache Cost: A Hands-on Measurement” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Reddit Codex: Diagnostic and Verification Prompt for Luna Subagent CompatibilityTurn “Reddit Codex: Diagnostic and Verification Prompt for Luna Subagent Compatibility” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.The Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and CachingTurn “The Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and Caching” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.Reddit Codex: Multi-Model Routing Configuration for Luna Subagents and Sol ReviewThis is not an official configuration, but a personal routing setup shared by a Codex user: use Luna for routine implementation, brainstorming, and subagents; use Sol high for plan reviews and final decisions; and use a fast model for test execution. It turns 。Use Cheap Luna to Orchestrate Threads: Use Threads Rather Than Same-Model SubagentsThe author recommends using Luna as a thread orchestrator, letting different threads choose Sol, Terra, or Luna for each task instead of having Sol Ultra automatically generate a batch of equally expensive subagents. The original prompt is very short. Its core。