TabbitBlog

GPT-6 Luna Pricing: API Rates, Effort Tiers, and Cost Calculator

GPT-6 Luna pricing explained: $0.10/$0.50 API rates, the 15x effort dial, cache and 272K rules, budget examples, and a local cost calculator.

In this article
  1. Key takeaways
  2. GPT-6 Luna API pricing at a glance
  3. The one number that decides your bill: the effort dial
  4. Translating the "50% cheaper" claim
  5. Cache economics: $0.01 reads, $0.125 writes
  6. The 272K long-context cliff
  7. Line items beyond the headline rate
  8. The family ladder and the open market
  9. What the rate card does not tell you
  10. A practical option: Tabbit Browser
  11. Two worked budget examples
  12. GPT-6 Luna pricing: which option should you choose?
  13. Verdict
  14. Sources

GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens on OpenAI's standard API rate card, checked September 23, 2026. Cached input is $0.01, cache writes are $0.125, and prompts above 272K input tokens reprice the whole request at long-context rates. The launch-day reaction on r/codex read the card in one line — "Sol now has terra pricing, Luna's price halved and terra is dead." — and that is the correct headline: against the GPT-5.6 Luna rates shown in the launch post ($0.20/$1.20), input fell 50% and output fell 58%. (r/codex thread; OpenAI pricing)

Reddit comment by PairStrong reading Sol now has terra pricing, Luna's price halved and terra is dead
r/codex, launch day: the top comment (172 points) on the price-card reaction thread; a community post, not an official statement.

But the sticker price is the least interesting number on this card. At $0.10/$0.50, the rate barely registers on most invoices — what decides your bill is which reasoning-effort tier you run (a 15x spread on identical token prices), your cache-hit share, and whether a prompt crosses 272K tokens. This page separates the API rate card from ChatGPT and Codex subscriptions, shows the full rate table with sources and dates, and ends with worked budgets and a calculator that runs in your browser. If your next step after budgeting is actually running Luna against web pages and files, Tabbit Browser is one client option; it does not change OpenAI's billing.

Key takeaways

  • Standard short-context GPT-6 Luna costs $0.10 input, $0.01 cached input, $0.125 cache writes, and $0.50 output per million tokens; an OpenAI spokesperson confirmed to VentureBeat these are permanent rates, not promotional.

  • The effort dial dominates the bill: Artificial Analysis measures the same card between $0.0045 per Index task (low) and $0.07 (max) — 15x — and the API default is medium, not max.

  • Above 272K input tokens, every token in the request moves to long-context rates ($0.20 / $0.02 / $0.25 / $0.75 per million).

  • Batch and Flex are listed at half of Standard; Fast is 2x; regional processing adds 10% where eligible, and Luna's EU data residency is Standard-only.

  • A token rate is not a subscription quota and not a per-task price: output volume (+24% vs the predecessor on AA's Index tasks), retries, and quality regressions all re-enter the budget.

GPT-6 Luna API pricing at a glance

All prices are USD per one million tokens, from OpenAI's pricing page and the gpt-6-luna model page, checked September 23, 2026. The long-context rows apply when input is more than 272,000 tokens.

GPT-6 Luna service tierContextInputCached inputCache writesOutput
Standard≤272K input$0.10$0.01$0.125$0.50
Standard>272K input$0.20$0.02$0.25$0.75
Batch / Flex≤272K input$0.05$0.005$0.0625$0.25
Batch / Flex>272K input$0.10$0.01$0.125$0.375
Fast≤272K input$0.20$0.02$0.25$1.00
Fast>272K input$0.40$0.04$0.50$1.50
OpenAI's flagship API pricing table showing gpt-6-astra, gpt-6-sol and gpt-6-luna rates for short and long context
OpenAI's flagship rate card as displayed on September 23, 2026; Luna is the bottom row, with short-context and long-context columns side by side.

The rate card sits inside a family ladder. These are Standard short-context rates from the same table:

ModelInputCached inputCache writesOutputWhat the row is useful for
GPT-6 Luna$0.10$0.01$0.125$0.50High-volume, tightly scoped work — the launch post's own positioning
GPT-6 Sol$2.00$0.20$2.50$10.00Repeated complex coding and agent workloads
GPT-6 Astra$10.00$1.00$12.50$50.00The flagship tier for the hardest tasks
GPT-5.6 Sol (predecessor, promotional)$4.00$0.40$5.00$20.00The comparison baseline OpenAI still lists, promo through at least Nov 21, 2026

Luna is twenty times cheaper than Sol on input and output, and a hundred times cheaper than Astra. For the sibling economics, read the GPT-6 Sol pricing breakdown and the GPT-6 Astra pricing breakdown; for what the cheap tier actually buys in capability — and where it regresses — read the GPT-6 Luna review. The predecessor's own price history is in the GPT-5.6 Sol cost guide.

The one number that decides your bill: the effort dial

gpt-6-luna accepts reasoning_effort values of none, low, medium (the default), high, xhigh, and max. The token prices do not change between them — the number of reasoning and output tokens does. Artificial Analysis tracks Luna as six entries, one per tier, and their measured economics (snapshotted September 23, 2026) look like this:

Reasoning effortIntelligence IndexCost per AA Index taskOutput speed
low21$0.0045175 t/s
Non-reasoning (none)18$0.01141 t/s
medium (default)29$0.02143 t/s
high32$0.03149 t/s
xhigh34$0.04153 t/s
max37$0.07157 t/s

Two details in this table matter more than the rates above:

  1. The spread is 15x. The card that looks like "$0.10/$0.50" bills between roughly half a cent and seven cents per AA Index task depending entirely on a request parameter. The default is medium ($0.02, Index 29). The headline intelligence (Index 37) exists only at max ($0.07). If you copy vendor benchmark numbers into a business case, budget for the tier the benchmark ran at.

  2. Low effort is cheaper than no reasoning. At $0.0045, low undercuts none at $0.01 on AA's workload, because non-reasoning runs emitted more output tokens per task. "Turn off thinking to save money" is not automatically the cheapest setting on this model — it is a hypothesis to test on your own traces.

OpenAI's own launch numbers point the same direction. On AutomationBench — a 47-tool business-workflow test — GPT-6 Luna at high effort improved on its predecessor by 5.4 percentage points at 58% lower cost per task; on internal factuality, "at higher effort levels it matches GPT-5.6 Sol at about a hundredth its cost." Read those as effort-conditional claims: the cheap tier reaches them when you pay for the thinking.

One launch-day user described the dial's practical effect on output shape: "5.6 Luna yapped a lot about unnecessary stuff, while 6 Luna is much more to the point… 6 Luna just gives you the solution and stops talking." Conciseness is part of the cost story — but AA's inspectors also tied knowledge-work regressions (skipped rubric sections, thinner deliverables) to the same token thrift, so treat terse output as a trade-off, not a free win. (r/codex comment)

Reddit comment by BelatedCube182 saying 6 Luna is much more to the point than 5.6 Luna
r/codex: a user comparing the two Lunas' verbosity on a Blender question; personal impression, not a measurement.

Translating the "50% cheaper" claim

The launch post's pricing card shows Luna moving from $0.20 → $0.10 input and $1.20 → $0.50 output, both labeled "50% cheaper." The arithmetic on output is −58.3%; OpenAI's framing counts whole-dollar price steps. Neither reading is wrong, but the baseline matters: those GPT-5.6 figures were themselves promotional rates, and VentureBeat reports a spokesperson confirming the new GPT-6 Sol and Luna prices are permanent, not promotional. (launch post; VentureBeat)

Independent measurement agrees on the direction and adds the fine print. Artificial Analysis measured Luna (max) at $0.07 per Index task versus $0.18 for GPT-5.6 Luna (max) — about 60% less — while the new model used more output tokens per task (51k vs 41k). The price cut carried the savings; token efficiency did not. Readers watching the launch charts caught the same shape: "it really looks like luna 6.0 is just as smart as luna 5.6 at less than half the cost." (r/codex benchmarks thread)

Reddit post in r/codex titled GPT 6-Luna benchmarks asking about Terminal Bench 4.0 and noting speed-for-cost complaints
r/codex, launch day: the benchmarks thread where users interpreted the AA chart; community discussion, not a measurement.

The wider framing some practitioners used is a price war. "GPT-6 Luna cut its token price by 90% in two months. July: $1/$6 per million input/output tokens. September: $0.10/$0.50… This is what an inference price war looks like," one AI engineer posted. The July figure — GPT-5.6 Luna's earlier standard rate, before its own promotional cut — is the poster's claim and can no longer be verified on OpenAI's current pages, so treat it as community context, not an official rate history. (X post)

X post by Mikel Echeve stating GPT-6 Luna cut its token price by 90 percent in two months
X, launch day: the price-war framing; a practitioner post, and its July baseline is not verifiable on current official pages.

Cache economics: $0.01 reads, $0.125 writes

For agent loops that re-send the same context, the cache rows matter more than the headline rates. Luna's cached input is $0.01 per million — a tenth of fresh input, the maximum 90% discount OpenAI advertises for GPT-6 caching — while cache writes are $0.125, a 25% premium over the input rate. Two accounting rules follow: write volumes are a separate line from reads, and a token should be counted once — as a write when the prefix is stored, and at the cached rate on turns that reuse it inside the 30-minute eligibility window OpenAI describes. (caching post; pricing page)

GPT-6 also changed caching behavior, not just prices: hit rates are higher by default, and new controls exist precisely because agents re-read context every turn — a caching dashboard, cache-miss diagnostics, explicit breakpoints, prewarming, and the ability to change reasoning effort between responses without breaking the cache (useful when you want max effort on the hard turn and low on the routine follow-up). OpenAI's post credits GitHub with cutting the share of prompt tokens requiring fresh processing by more than 50% across billions of requests, and one customer with a 20% cost cut from a few points of extra hit rate.

Practitioners noticed the economics immediately. "Luna is my favorite model for building product features thanks to its cost (and speed)," wrote Simon Willison on launch day — the routing pattern that makes $0.01 cached reads and 141–175 t/s generation the working combination for product surfaces. (X post)

X post by Simon Willison calling GPT-6 Luna his favorite model for building product features thanks to cost and speed
X, launch day: a widely-followed practitioner on Luna's cost/speed position, quoting the OpenAI announcement; personal preference, not a benchmark.

The 272K long-context cliff

The model card's sentence is short: prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request. This is a request-level threshold, not a surcharge on the excess tokens. A 273,000-token prompt does not pay short-context rates on its first 272K — every billed category, including output, uses the long-context row.

The window itself is 1,050,000 tokens with 128,000 max output. "Fits in context" and "same price per token" are different claims. If your prompts hover near the threshold, keep margin: one appended document can reprice the entire request. This is the same cliff structure as Sol's — the GPT-6 Sol pricing guide works through it in detail, and the same discipline applies here at one-tenth the scale.

Line items beyond the headline rate

Batch and Flex halve the card; Fast doubles it. OpenAI lists Batch and Flex at 50% of Standard rates — the largest flat lever for latency-tolerant work. Fast mode (Priority was renamed to Fast on July 30, 2026) is 2x. Verify your endpoint actually uses the tier before booking either multiplier.

Regional processing adds 10%. Eligible data-residency endpoints for models released on or after March 5, 2026 carry a 10% uplift, and GPT-6 Luna's EU data residency is available only with Standard processing.

Tool tokens bill at model rates. Tokens consumed by built-in tools are billed at the selected model's token rates; some tools and hosted sessions carry separate fees. Failed attempts that consumed tokens belong in the budget too.

A "Pro" variant exists on routers. OpenRouter lists "GPT-6 Luna Pro" — the same underlying model, "served with reasoning.mode set to pro for higher-quality responses on complex tasks" — at the same $0.10/$0.50 base price. It is a provider-side serving configuration, not an additional OpenAI rate-card row; treat it as a routing option to evaluate, not a different model. (OpenRouter)

Rate limits: the free tier is closed. The model page marks gpt-6-luna "Not supported" on Free, with paid access from Tier 1 (500 RPM / 500K TPM / 5M batch-queue tokens) through Tier 5 (30,000 RPM / 180M TPM). Notably, OpenAI publishes higher token-per-minute limits for Luna than for Sol at equal paid tiers — Tier 2 is 2M TPM for Luna versus 1M for Sol — which reads as an invitation to use the cheap model at volume. (model page)

API tierRPMTPMBatch queue limit
FreeNot supported
Tier 1500500,0005M tokens
Tier 25,0002,000,00020M tokens
Tier 35,0004,000,00040M tokens

Source: OpenAI's gpt-6-luna model page, checked September 23, 2026.

The family ladder and the open market

Inside OpenAI's ladder, the routing logic is straightforward: try Luna first for tightly scoped, high-volume steps; escalate to Sol when a failed attempt is expensive; reserve Astra for the hardest work. The launch post itself positions Luna for "more tightly defined jobs such as summarization, extraction and answering straightforward questions."

The counterintuitive finding is that Luna is not the cheapest capable API on the open market. VentureBeat's launch-day comparison table lists Meta's Muse Spark 1.2/1.3 Contributor at $0.10/$0.20 and Xiaomi's MiMo-V2.6-Flash at $0.14/$0.28 — both with lower combined input+output rates than Luna's $0.60 — and DeepSeek's V4.1 Flash at $0.15/$0.60 off-peak. What Luna buys with its premium is the OpenAI ecosystem: a 1.05M-token window, the GPT-6 tooling and caching stack, and effort-tier headroom from $0.0045 tasks upward. If raw floor price is the only criterion, the open market undercuts it. (VentureBeat table)

Budget modelInput / Output per 1MCombinedNotable constraint
Muse Spark 1.2/1.3 (Contributor)$0.10 / $0.20$0.30Contributor-tier availability
MiMo-V2.6-Flash$0.14 / $0.28$0.42Smaller ecosystem
GPT-6 Luna$0.10 / $0.50$0.60Free API tier closed; format-trade-offs (see review)
DeepSeek-V4.1-Flash (off-peak)$0.15 / $0.60$0.75Peak rates double; text-only

The community's read on the trajectory is expansive: "I feel 2027 will be the year when intelligence will be economical enough to use at most of the places," the most-upvoted reply in r/OpenAI's launch thread predicted. Useful direction, unverifiable as a forecast. (r/OpenAI thread)

Reddit post in r/OpenAI titled Intelligence Too Cheap To Meter finally being true with GPT 6 Sol and Luna
r/OpenAI, launch day (146 points): the excitement pole of the pricing reaction; community post.

What the rate card does not tell you

It is not a subscription calculator. ChatGPT and Codex plans set their own usage limits; a 50% token-price cut does not mathematically double any plan's allowance. At launch, GPT-6 Luna reached ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, while Free and Go plans got it only in the desktop app. The subscription experience surfaced its own complaints: "it's faster on xhigh and my usage is barely going down," one Codex user reported, and in a pricing thread the terse summary was "Reduced pricing and reduced limits." API arithmetic and plan quotas are different purchases — check the live usage page for your plan. (usage comment; pricing thread)

Reddit comment by Razbari saying GPT-6 Luna is faster on xhigh and their usage is barely going down
r/codex: a subscriber's observation that cheaper API rates did not visibly move plan usage; personal account, not an official limit.

It is not a per-task price either. AA measured Luna using 24% more output tokens per Index task than its predecessor, and its inspectors attributed knowledge-work regressions (~75 Elo on GDPval-AA, ~45 on AA-Briefcase) to skipped presentation and rubric elements. A task that Luna fails or truncates gets retried — on a costlier model if you escalate — and retries are billed. The GPT-6 Luna review covers the quality evidence; the budgeting consequence is that cost per successful task, not cost per token, is the number to track: total spend including failures, divided by tasks you actually shipped.

A practical option: Tabbit Browser

If the budget question is settled and the next question is where to run Luna against real work — open pages, files, research tasks — Tabbit Browser is one answer that does not require wiring an API integration. Its model picker lets you select models like GPT-6 Luna alongside your open tabs, and the GPT-6 Luna model hub collects the browser-side references: prompt notes for task-shaped requests and review notes for what users reported.

The honest boundary: Tabbit is a client. It does not change OpenAI's rates, and an API token estimate on this page does not predict any subscription's or client's usage. Model availability depends on the live picker for your account and edition. This article has no measured Tabbit bill, task-success rate, or latency result for Luna — if you need those numbers, run your own fixed task set and keep the receipts.

Tabbit Browser

For the workflow context — why an AI browser or an agentic browser changes where a cheap worker model is useful — those guides cover the client side; the rate card above remains the source for token billing.

Two worked budget examples

The examples use Standard short-context rates, no regional uplift, tool fees, retries, or tax. They are arithmetic illustrations, not measured workloads.

Short request. 10,000 uncached input tokens and 2,000 output tokens:

(10,000 × $0.10 + 2,000 × $0.50) / 1,000,000 = $0.002

Warm agent loop. A 40-turn run that re-sends a 250,000-token prompt every turn, with 90% of it served from cache (written once, reused inside the 30-minute window) and 3,000 output tokens per turn. Per turn:

(225,000 × $0.01 + 25,000 × $0.10 + 3,000 × $0.50) / 1,000,000 = $0.00625 per turn

...plus the one-time write of the stored 225K-token prefix at $0.125 per million (about $0.03). Across 40 turns the cached loop bills roughly $0.28, where re-reading the same prompt uncached every turn would cost about $1.06 — close to a 4x difference on identical token prices, decided entirely by the cache-hit share. At Luna's scale that is pennies either way; multiply by a million monthly turns and the cached pattern bills about $7,000 (writes included) versus roughly $26,500 cold. That is the shape of workload where Luna's $0.01 cached-input row, not its headline price, is the real news.

For planning: monthly estimate = per-task average × monthly task count; cost per successful task = total spend including failed attempts ÷ successful tasks. Subscription quotas and any client-side fees cannot be derived from these formulas.

Local calculator

Estimate GPT-6 Luna API cost

Official USD rates checked Sep 23, 2026. Values stay in this browser.

Per task$0.00200
Per month$0.00

Excludes tools, retries, taxes, and regional processing. Cache writes are separate from cached reads. Check current rates

The calculator runs in your browser and sends nothing anywhere. It estimates token charges from your input volume, cache-hit share, cache writes, output, run count, and service tier; it switches to long-context rates above 272K input. It does not include regional processing, tool fees, taxes, provider markups, retries, or subscription limits.

GPT-6 Luna pricing: which option should you choose?

If your priority is…Start with…Keep in the estimate
Highest volume, tightly scoped tasksLuna at low or medium effortEffort tier spread (15x), output-token growth, retries
Deep-but-cheap reasoning on hard stepsLuna at high/xhigh, cache the prefix30-minute window; effort changes keep the cache
Offline or batch pipelinesBatch or Flex where eligible50% tier rate; queue and latency constraints
Prompts near 272KTrim context or budget the long rowWhole-request repricing above the threshold
Interactive latencylow effort or none; Fast only if neededFast doubles rates; TTFT differs by tier
Consumer ChatGPT/Codex useThe relevant plan's own limitsPlan quotas are not API token math

Verdict

GPT-6 Luna's rate card is the cheapest OpenAI has ever sold a GPT-6-class model at: $0.10/$0.50 per million, permanent per the company's spokesperson, with cached input at a tenth of that. The honest budget treats the sticker as the floor, not the bill. The effort dial alone moves measured per-task cost 15x between low and max; output volume grew against the predecessor; the 272K cliff reprices whole requests; and quality regressions convert into retry costs at exactly the workloads where you would otherwise celebrate the price. Set effort deliberately, cache aggressively, watch cost per successful task — and compare subscriptions on their published allowances, never on token arithmetic.

Launch-day reactions are directional, not settled: excitement about "intelligence too cheap to meter," caution that "the consequence of its work will not be seen or measured… until a few days." OpenAI's pricing page is the live source; re-check it before committing a budget.

Sources

FAQ

How much does GPT-6 Luna cost per million tokens?

OpenAI's standard short-context API rates are $0.10 per million input tokens, $0.01 per million cached input tokens, $0.125 per million cache-write tokens, and $0.50 per million output tokens. Above 272,000 input tokens, the full request uses the long-context rates of $0.20, $0.02, $0.25, and $0.75 respectively.

Is GPT-6 Luna's API pricing permanent or promotional?

OpenAI's launch post frames the new rates as 50% below the GPT-5.6 promotional pricing, and a company spokesperson told VentureBeat the GPT-6 Sol and Luna rates are permanent prices, not promotional or introductory pricing. The predecessor GPT-5.6 Sol comparison rate on OpenAI's pricing page is still labeled promotional through at least November 21, 2026.

What is the cheapest way to run GPT-6 Luna?

Three levers dominate the bill: reasoning effort, caching, and service tier. Artificial Analysis measures the same rate card between $0.0045 per Index task at low effort and $0.07 at max, cached input reads cost a tenth of fresh input, and Batch or Flex service tiers are listed at 50% of Standard rates for work that tolerates their timing.

Does GPT-6 Luna have long-context pricing?

Yes. Prompts with more than 272,000 input tokens are priced at 2x the input and cache rates and 1.5x the output rate for the entire request, not only the tokens above the threshold. The context window itself is 1,050,000 tokens, so fitting in context does not mean paying the short-context rate.

Is GPT-6 Luna API pricing the same as ChatGPT or Codex usage?

No. The API bills tokens; ChatGPT and Codex plans define their own usage limits. At launch, GPT-6 Luna was available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, while Free and Go plans could only use it in the desktop app. A cheaper token rate does not automatically mean more subscription usage.

Can I use GPT-6 Luna on the free API tier?

No. The model page's rate-limit table marks gpt-6-luna "Not supported" on the Free tier, with paid access starting at Tier 1 (500 requests per minute, 500,000 tokens per minute). OpenAI also publishes higher token-per-minute limits for Luna than for Sol at the same paid tiers.

Take the next step

Let Tabbit work alongside you.

Research across tabs, automate repetitive browser work, and keep every piece of context within reach.