OpenAI shipped GPT-6 Sol and GPT-6 Luna on 22 September 2026, and the interesting part is not the benchmarks. Sol costs $2 per million input tokens and $10 per million output. Luna costs $0.10 and $0.50. The same family that charges $10 and $50 for Astra now has a tier that is one hundred times cheaper on input.
What actually launched
GPT-6 Astra arrived first, positioned as OpenAI's most capable and most aligned model. Sol and Luna followed on 22 September, built with the same methods and aimed at the price tiers Astra cannot serve. OpenAI's framing is that the Astra improvements now reach cheaper models; the press framing was "mini-Astras", which is close enough.
Both carry a 1,050,000-token context window and a 128,000-token maximum output, the same envelope as Astra. Both accept text, images and files. That is the part worth sitting with: the cheap tier is not a cut-down context or a text-only model. It is the same shape at a fraction of the cost.
One correction to the rumour mill, since it circulated widely: there is no GPT-6 Terra. The GPT-6 family is Astra, Sol and Luna. Terra exists in the 5.6 generation and did not carry forward.
The price is the release
Run the numbers on a realistic message. A 3,000-token prompt with a 1,500-token answer costs about $0.021 on Sol. The same message on Astra costs about $0.105, five times more. On Luna it is about $0.001.
That spread changes what you can build. Work you would previously have batched overnight because the flagship was too expensive now runs interactively on Sol. Work you would not have automated at all, because even the cheap tier cost too much at volume, is plausible on Luna.
It also changes what "use the best model" means. When the gap between tiers was small, defaulting to the flagship was defensible. At a five-times spread, it is a decision you should be able to justify per task.
Where each one earns its place
Sol
The all-rounder. OpenAI positions it for complex work including coding, and it is the tier most people should default to for anything where the answer matters. On Deepest, Sol is what Auto now selects when you do not pick a model yourself, for general, creative and vision tasks.
Luna
The volume tier. At $0.50 per million output tokens it is cheap enough that the cost stops being the thing you think about. Classification, extraction, summarisation, first-pass drafting, anything you run hundreds of times.
Astra
Still the ceiling, and still priced like it. Worth it when the task genuinely needs the top of the range and you have checked that it does.
The comparison that matters
The same week gave us Claude Opus 5.5 at $4 and $20, and Grok 4.7 at $2 and $6 on xAI's list. Sol at $2 and $10 sits between them. On input, Sol and Grok 4.7 are level and Opus 5.5 is double. On output, Grok is cheapest, Sol is middle, Opus is highest.
Price tells you almost nothing about which one answers your question better. Three frontier models landed within 48 hours of each other at prices close enough that cost is no longer the deciding factor. What decides it is which one is good at your particular kind of problem, and that is not something a benchmark table answers for you.
This is the argument for asking more than one. On Deepest you can send the same prompt to Sol, Opus 5.5 and Grok 4.7 together and read the three answers side by side. When they agree, you can move on. When they disagree, the disagreement is the useful part, and it is invisible if you only ever asked one.
Frequently Asked Questions
Is GPT-6 Luna good enough to use as a default?
For high-volume, well-specified work, yes. For anything open-ended or where being wrong is expensive, start with Sol. The cost difference is real but so is the capability difference, and a cheap wrong answer is not a saving.
Does GPT-6 replace GPT-5.6 entirely?
Not yet. GPT-5.6 Sol, Terra and their pro variants are still available and still capable. On Deepest we retired the 5.6 Luna pair when GPT-6 Luna arrived at under half the price, since there was no case left for the older one, and anyone who had it selected now resolves to the newer model automatically.
How do I tell which of these is best for my work?
Run your own prompts through several of them rather than reading benchmark tables. Benchmarks measure average performance on standardised tasks; you care about your tasks. Deepest exists largely because that comparison is tedious to do by hand across three separate subscriptions.