All articles
AI News

GPT-5.6 Sol vs Terra vs Luna: Which Tier Should You Use?

OpenAI shipped all three tiers with pro variants on 9 July. A five-times price spread from Luna to Sol, identical context and output limits, and how to choose between them.

Travis Johnson

Travis Johnson

Founder, Deepest

July 13, 20263 min read

OpenAI released six models on 9 July: GPT-5.6 Sol, Terra and Luna, each with a pro variant. Not a flagship with a cheap sibling bolted on later, but a complete price ladder shipped in a single announcement. That is a change in how these launches work, and it puts the choice of tier squarely on you.

The ladder

The three tiers run Luna, Terra, Sol from least to most capable. Launch pricing per million tokens is $1 in and $6 out for Luna, $2.50 and $15 for Terra, and $5 and $30 for Sol. Each has a pro variant that runs the same model with more reasoning effort.

All six share the same envelope: a 1,050,000-token context window and a 128,000-token maximum output, accepting text, images and files. The tiers differ in capability and price, not in what they will accept or how much they will write.

OpenAI positions Sol for enterprise work, coding, scientific research and cybersecurity, describing it as their strongest cybersecurity model to date. The models were previewed to a small group of partners in late June before the public release.

Key Finding: The gap from Luna to Sol is five times on input and five times on output. That is a real decision, not a rounding error. Defaulting every request to the top tier out of habit is now a choice you are making about your budget, whether or not you have noticed you are making it.

Speed is becoming a product dimension

Alongside the launch, OpenAI made Sol available on Cerebras hardware at up to 750 tokens per second. That is fast enough to change the character of an interaction rather than just shortening the wait.

It is worth separating two things that often get conflated. Latency to first token determines whether something feels responsive. Throughput determines how long a long answer takes to finish. A model generating at 750 tokens per second finishes a 2,000-token answer in under three seconds, which is the difference between reading along and waiting for a wall of text to land.

Expect this to become a standard axis of competition. When several models are close on quality, the one that finishes first wins a lot of everyday work.

What the pro variants are for

The pro variants are not different models. They are the same weights configured to spend more reasoning effort before answering. That extra thinking is billed as output tokens, so a pro variant costs more per useful word even where the headline rate is identical.

Our advice: start on the standard variant. Move to pro when you have a specific class of problem where the standard tier visibly fails, and check that pro actually fixes it rather than assuming more thinking helps. On plenty of tasks it does not, and you have simply paid for a longer internal monologue.

How to choose a tier without guessing

The instinct is to pick Sol for important work and Luna for everything else. That is a reasonable starting heuristic and a poor stopping point, because "important" is not the same as "hard", and a five-times price gap deserves better than instinct.

The approach that actually works is boring: take ten prompts representative of your real work, run them through all three tiers, and read the answers side by side. You will usually find a clear line. Some categories genuinely need Sol. Rather more of them do not, and the difference is invisible until you look at the outputs next to each other.

This is the comparison Deepest is built for. Send one prompt to Sol, Terra and Luna at once and the three answers arrive together, which turns a tedious afternoon of manual testing into a single message.

Frequently Asked Questions

Which GPT-5.6 tier should I default to?

Terra is a sensible default for mixed work: it sits at half Sol's price and well above Luna's capability. Move up for genuinely hard reasoning and down for high-volume, well-specified tasks. The important thing is to make the choice deliberately rather than inheriting whatever the interface picked for you.

Do the pro variants cost more per token?

At launch they carry the same headline rate as their standard counterparts, but they spend more output tokens on reasoning before answering. Since reasoning is billed as output, the effective cost of a pro answer is higher even where the rate card looks identical.

Is a 1M context window actually useful?

Less often than it sounds. Very long contexts cost real money to send and models do not attend uniformly across them, so retrieval that finds the right 10,000 tokens usually beats dumping a million. The large window is best treated as headroom for the occasional case, not a default working style.

GPT-5.6OpenAImodel pricingAI newsGPT-5.6 SolGPT-5.6 LunaCerebras

See it for yourself

Run any prompt across ChatGPT, Claude, Gemini, and 300+ other models simultaneously.30-day money-back guarantee.

Try Deepest free →

Related articles