Skip to content
AiHub
← Writing

Reliability

What happens when a model fails, and who pays for it

Providers go down, rate-limit, time out and sometimes return nothing at all. What each of those looks like from inside the app, which ones touch your balance, and the one thing we will not do to hide them.

· 7 min read

Every AI product is a thin layer over somebody else’s servers. The interesting question is not whether those servers fail — they do, most weeks — but what the app does in the ten seconds afterwards, and whose money it does it with.

Five different things people call “it broke”

From the outside a failure is one event: you asked something and did not get an answer. From inside the gateway there are five, and they have almost nothing in common except the face they show you.

  • It never left. Your balance was too low, or the request was malformed — an option a model does not accept, a file it cannot read.
  • The provider refused. A rate limit, an outage, an endpoint that answers every request with a shrug.
  • The provider accepted and returned nothing. A 200 with an empty body, which is the most confusing failure there is.
  • It stopped halfway. Three paragraphs arrived and then the connection died.
  • It was simply slow. Not a failure at all, but indistinguishable from one if whoever is waiting gives up first.

The failures that cost you nothing

Three of those five are free, and not by policy — by construction. The charge is written from the usage the provider reports once the work is done, so a request that never reached a provider has nothing to write a charge from. There is no estimate standing in for it and no minimum fee.

The low-balance case is the clearest. An account that cannot afford the request is turned away before anything is sent upstream, with the balance in the refusal so the app can say what is actually wrong. Nobody should be billed for being told no, and the only reliable way to guarantee that is to check first.

The empty-response case is stranger and worth stating plainly: if a provider takes the request, returns 200, and hands back nothing usable, you are charged nothing. We could compute something from the length of what you sent. We would rather absorb it than invent a number and put it in your ledger, where it would sit forever looking exactly like a real charge.

The awkward one: a reply that stops halfway

Streaming is why chat feels alive, and it is also why chat can fail in the middle of a sentence. When that happens, three things are true at once: you have some of an answer, the provider did some real work, and neither of us can get that work back.

So the partial text stays in the transcript rather than vanishing — you may well keep it — and the charge is whatever the provider reports for the part that ran, doubled like everything else. Usually that is a fraction of a whole reply. The failure is drawn in place of the rest of the message, with a retry underneath it, rather than as a banner somewhere else on the screen: the thing that broke is the thing that should be carrying the explanation.

One detail took a second pass to get right. Provider errors arrive as English sentences, and for a while an app running in Persian would speak Persian right up until the moment it had bad news, then quote a stranger verbatim. Upstream failures are now said in your own language, like every other message in the interface.

When a whole provider goes dark

The most instructive failure we have had was not intermittent. Our main provider’s Claude endpoint began answering every request with “Server exception”, and kept doing it. We reproduced it by calling them directly with their own documentation open, which is the only way to be sure a fault is theirs and not a misreading of it in your own client.

All ten Claude models now answer through a second route to the same models, charged at the cost that route actually reports, doubled by the same rule as everything else. From inside the app nothing about that is visible, which is the point — the promise on a model card is the model, not the road it travels down.

The honest footnote is that this only works where a second road exists. It does for the big text models and it does not for everything, and when a model is genuinely unreachable the app says so instead of offering you a different one.

The free assistant is a chain for exactly this reason

AiHub AI, the assistant that costs nothing, is not one model behind the scenes. It is a short chain: several models that providers publish at no cost, and one cheap paid model at the end that we carry ourselves. Free capacity is the least reliable capacity on the internet — it rate-limits at busy hours, it goes down, it gets retired without an announcement.

When the first link refuses, the request walks to the next one inside the same call. You see an answer arrive a moment later than it might have; you do not see an error, and you do not see a message asking you to try again. A free tier that breaks is not a free tier, it is an advert for the paid one, and that is not what ours is for.

Slow is not broken

The failure that fooled us hardest was not a failure. During a sweep of every chat model in the catalogue, one deep reasoning model came back empty and got written down as broken. It was not: it had been thinking, and our own test client stopped listening at seventy seconds. Given room, it answered perfectly.

This is a real trap in a product where the price of a model correlates with how long it takes. The models worth the most money are exactly the ones most likely to be killed by an impatient timeout — and a timeout on our side, after the provider has done the work, is the worst of both worlds: you wait, we pay, nobody gets an answer. A model that thinks for a minute now gets its minute.

Pictures and video fail differently

A generation is not a conversation. It is a job: you agree to a quoted price, the provider queues it, and one to five minutes later there is a file or there is not.

Because the charge comes from the provider’s own report, a job that dies produces no report and therefore no charge — the quote you agreed to is a ceiling, not a deposit. What a failed generation actually costs you is the wait, which is why the advice in the piece on pictures and video is to draft at the cheap settings. Five minutes lost on a 480p attempt is an annoyance; five minutes lost at 1080p is the same annoyance and you are no closer.

The failure that cost us instead of you

An adversarial pass over the whole product last month found a hole worth admitting to, because it went the direction these things usually do not. A chat request could be delivered and never charged: the check before the call used a flat estimate while the provider bills for the whole transcript, and if the charge afterwards failed it was written to a log and forgotten. The reply went out free, and the same request could be repeated indefinitely.

Both halves are closed. We are telling you about it because a page that only lists the failures where the customer is protected is marketing. The reason to trust the ledger is not that we say it is right — it is that every row carries the provider’s own reported cost underneath what you paid, so you can check the arithmetic without asking us.

The one thing we will not do

There is a tempting fix for all of this, and it is common: when the model somebody chose is unavailable, quietly answer with a cheaper one. It removes the error, it removes the support email, and almost nobody notices.

We do not do it. If you picked a model, the answer comes from that model or it does not come at all, and you are told which. The single exception is the free assistant, which is a chain by design and says so on its own card. An answer from a model you did not choose is not a recovered failure. It is a silent one, and those are the expensive kind — you find out weeks later, in work you have already sent to somebody else.

What a failure costs, stage by stage, is written out on the refunds page, and the ledger behind your balance shows the provider’s figure under every charge.