Pricing
What a model actually costs, and why we double it
Providers bill in fractions of a credit for work that has already happened. Here is the whole chain from their invoice to your balance, and why the multiplier is printed rather than buried.
· 5 min read
Every AI product marks up the model. Almost none of them will tell you by how much. Here is the whole chain, from the provider’s invoice to your balance.
Nobody bills in words
The first thing to understand about AI pricing is that the unit you see is almost never the unit anyone is billed in. A model provider charges for tokens of text in and tokens of text out, at different rates, plus a separate rate for an image attached to the message, plus a different scheme entirely for pictures and video — usually a flat price per generation that changes with resolution and length.
A product built on top of that has to pick something to show you. The honest options are to pass the provider’s unit straight through, or to convert it into one unit and say exactly what the conversion is. The dishonest options are more popular: a monthly fee with a fair use policy nobody defines, or credits whose relationship to the underlying cost is deliberately unstated so it can be changed later.
The chain, end to end
On AiHub the chain has four links and no hidden steps.
- The model runs. Your request goes to our gateway, which sends it to the provider. The apps never hold a provider key, so there is exactly one place where cost is known.
- The provider reports what it consumed. Not an estimate from us — a number in their response, after the work is done, in their own credit unit.
- We double it. One multiplication, in a database transaction that locks your balance so two simultaneous requests cannot overdraw it.
- Both numbers are written down. The ledger row keeps the provider’s reported cost next to what you paid. You can divide one by the other. It is always two.
A token on your balance is therefore not an invented currency. It is one provider credit, and you are charged two of them for every one a request actually used. That is the entire conversion, and it is why the pricing page can show both columns without embarrassment.
Measured, not estimated
There is a shortcut nearly everyone takes here, and it is worth knowing about because it costs you money. It is much easier to estimate a request’s cost from its inputs — count the characters, apply the published rate, charge that — than to wait for the provider to say what happened. Estimating is faster, it lets you charge before the work finishes, and it is always a little bit generous to whoever wrote the estimator.
We do not do it. The charge is applied after the response arrives, from the reported figure. The practical consequence: if a model turns out to be cheaper than expected on your particular request, you get the cheaper price. Nobody has to notice or complain for that to happen.
For pictures and video, which are quoted before they run, we also keep a table of real generations we have already paid for — model, resolution, length, and the actual reported cost. When your settings match a row in that table, the quote is that recorded price. When they do not, the app says the number is an estimate rather than printing it as though somebody measured it. That distinction is on screen, because a confident wrong number is worse than an honest uncertain one.
What a subscription is actually selling you
A flat monthly fee looks simpler and usually is not. Underneath it, the maths only works if light users subsidise heavy ones — which means the provider needs light users, and needs heavy ones to stop. That is where the invisible machinery comes from: rate limits that tighten near the end of the month, a quiet switch to a smaller model under load, a fair use policy with no numbers in it, a queue that gets slower the more you use it.
None of that is villainy. It is arithmetic. But it means the price you agreed to is not the price you experience, and you have no way to audit the difference.
Per-request pricing has the opposite property. We have no reason to slow you down, because a busy account is a good account rather than a loss. And you can work out your own bill in advance, which is the thing a subscription can never let you do.
Why two, and not 1.4 or 3
Two is the smallest multiple that covers what sits between you and the model: storage for your history and your files, the gateway that routes around providers when they fail, the price table that only exists because somebody paid to run those generations, and the apps themselves. It is also a number you can do in your head, which matters more than it sounds. A 1.37× markup is not more honest for being precise — it is just harder to check.
The rule is the same on the cheapest chat model and the most expensive video model. There is no tier where it improves, and there is no model where it quietly worsens. If we ever change it, that is an email to every account before it takes effect, not a line edit to a page.
What it looks like on a real bill
The spread between the cheapest and dearest thing in the catalogue is larger than most people expect, and it is the single most useful fact for controlling what you spend.
- A message to Claude Haiku 4.5: about 0.06 tokens — three hundredths of a credit, doubled.
- The same message to Claude Opus 4.8: about 0.70 tokens. Roughly twelve times the price for the model that thinks hardest.
- A picture from Nano Banana 2 at 1K: 16 tokens. Around 260 Haiku messages.
- Four seconds of Seedance 2.5 at 480p: 224 tokens. Around 3,700 Haiku messages, for four seconds of video.
Chat is essentially free compared to pixels. If your bill is a surprise, it is almost never the conversations — it is one afternoon of video experiments. That is also why the video models quote you first and wait.
The part we cannot fix
We do not set the underlying prices, and providers change them. When a model gets cheaper upstream, your price falls the same day, because the charge is derived from their number rather than from a table we maintain. When one gets dearer, it rises the same way. We would rather that than a fixed price list that is quietly wrong for a month, and we are not going to pretend to a stability we do not control.
What we can promise is the multiple, the receipt underneath every charge, and the quote before anything expensive runs. Everything else on the pricing page is just the current numbers.