Deepest used to charge a flat number of credits per model per message. Claude Opus cost 30 credits whether you asked it the time or handed it a 200-page contract. That was simple, and it was wrong in both directions at once. Credits are now metered on what a message actually consumes.
Why flat pricing had to go
A flat per-message price has to cover the worst case, because an unmetered message might produce one. Our ceiling was an 8,000-token answer, so every model's flat price was set high enough to survive that.
Then we measured what messages actually look like. Across 470 real calls the median answer was 1,516 tokens, under a fifth of the ceiling. So the overwhelming majority of messages were being charged for an outcome they never came close to reaching.
Meanwhile the rare message that did hit the ceiling lost us money anyway. One message in September carried 181,475 tokens of attached documents. It cost us $2.21 to serve and billed 100 credits, worth about $0.40. Flat pricing managed to overcharge the quiet majority and undercharge the expensive minority simultaneously.
How it works now
Every message is charged on the tokens it actually uses: what you send, including anything you attach, plus the length of the answer. One rate applies across every plan, so credits mean the same thing whether you are on the cheapest tier or the largest.
Ask a flagship model a one-line question and you pay for a one-line question. Hand the same model a long document and you pay for the document. The model you pick still matters, because the rates differ by an order of magnitude between tiers, but it is no longer the only thing that matters.
You see the number before you send it
Metering only works for you if the cost is visible before you commit to it. Otherwise it just replaces a predictable price with an unpredictable bill, which is worse.
So the composer shows the estimate for the message you have actually written, with your attachments counted in, before you send it. The model picker shows a typical cost per model so you can compare while choosing. Both are marked as estimates, because the true figure depends on how long the answer turns out to be.
Attachments are deliberately quoted at the maximum they are allowed to consume rather than their real extracted length, which nothing can know until the file is processed. That means a small PDF is over-quoted and the final charge comes in under the estimate. We would rather be wrong in that direction.
What this means per plan
Every tier sells credits at the same rate. That is a change: we used to discount the top tier heavily, which sounds generous and in practice lost money on exactly the customers using the most. The cost of serving a credit does not fall with volume, so pricing as though it does is a slow way to go broke.
Unused monthly credits still roll over, capped at one full month's allowance. Pack credits do not expire.
The part we got wrong
When we first modelled this we concluded metering would be cheaper for typical users and nearly said so publicly. Then we measured properly across the full sample and found that just over half of real messages cost more, not less. The pleasant version of the story was not true.
We mention it because pricing pages tend to describe changes exclusively as wins for the customer, and that is rarely how pricing changes work. This one makes short messages cheaper, makes long ones honest, and moves some cost onto heavy users who were previously subsidised by everyone else. If you send long documents to expensive models, you will pay more than you used to. That is the intended behaviour, not a side effect.
Frequently Asked Questions
How do I keep my costs down?
Two things dominate: the model you choose and the length of what you send. The spread between tiers is roughly a hundred to one on input, so moving routine work to a cheaper model saves far more than trimming prompts. Attach documents only when the model needs them, since a large attachment can cost more than the rest of the message put together.
Why not just show prices in dollars?
Because you would then be reading four-decimal figures for every message, and the numbers move whenever a provider changes a rate. Credits give a stable unit that we absorb rate changes behind. The tradeoff is a layer of indirection, which is why the cost of your specific message is shown before you send it rather than buried on a pricing page.
Does the estimate ever come in under the real charge?
It can, because the answer's length is genuinely unknown until the model writes it, and the estimate assumes a median-length reply. An unusually long answer can exceed the estimate. Attachments, which are the largest and most surprising cost, are deliberately quoted high so that side errs the other way.