Skip to content
AiHub
← Writing

Choosing a model

Pick a model by the job, not by the brand

Brand loyalty in AI is expensive. Which differences between models are real, which are leaderboard noise, and when the model that thinks hardest is worth twelve times the price.

· 5 min read

Most people pick a model the way they pick a brand of coffee: once, on reputation, and then never again. It is the most expensive habit in AI.

The badge is not the job

The models are not general-purpose in the way the marketing suggests. They are general-purpose the way a Swiss Army knife is: everything is in there, and three of the tools are much better than the rest. The frontier model that writes beautiful prose is twelve times the price of the small one that can look something up, and it is not twelve times better at looking something up. It is the same.

This is why AiHub groups models by what you want done rather than by who made them. Nobody sits down wanting to use a company. They sit down wanting to think something through, get a fast answer, write something, check a fact, read a long document, or make a picture.

The question that decides it

Before picking, answer one thing: what would a bad answer cost me?

If a bad answer costs you fifteen seconds — a definition, a synonym, a shell flag, a quick reformat — use the cheapest thing on the shelf. Claude Haiku 4.5 and Gemini 2.5 Flash are both around 0.06 tokens a message. You could ask them a hundred questions for less than the price of one small image, and if one is wrong you will know immediately.

If a bad answer costs you an hour of rework, or goes into something someone else reads, that is when the expensive models earn their keep. Claude Opus 4.8 at about 0.70 a message is twelve Haikus. It is worth it for the piece of work where you cannot easily tell that the answer is wrong.

The mistake is not using the expensive model. It is using it for the fifteen-second questions, all day, out of habit.

Where the differences are actually real

Benchmark tables move a percent at a time and change every month. These differences do not, and they are the ones worth choosing on.

  • Long documents. Context length is a hard limit, not a matter of taste. A model that cannot hold your 200-page PDF will do a confident job on the part it saw and say nothing about the rest. The Gemini Pro models are the ones built for this, and they are on the “read a long document” shelf for that reason.
  • Attachments at all. Some models simply cannot take a file. In our catalogue the Claude family, the Gemini family and GPT 5.2 accept them; the Codex coding variants and both Grok models do not. If your question is really “what does this screenshot say”, that single fact eliminates half the list before you think about quality.
  • Prose. There is a real difference between a model that writes correct sentences and one that writes sentences with a rhythm. Claude Fable 5 exists for the second thing. If you are drafting anything a person will read for pleasure, it is worth the 0.30.
  • Current information. A model’s training has an end date, and past it, confident invention is the default failure. Grok 4.5 and the grounded Codex variant sit on the “keep it factual” shelf because they are the ones that go and look.
  • Code. The Codex family is tuned for it, at three prices: fast, grounded, and deeper. The deep one is 0.24 a message; the fast one is 0.12. For “why does this stack trace say that”, take the fast one.

Start cheap and escalate

The routine that saves the most money is embarrassingly simple: ask the cheap model first. Most of the time the answer is fine and you are done for 0.06. When it is not — when the answer is vague, or wrong in a way you can spot, or it misses the point of the question — switch model in the same conversation and ask again. The thread comes with you; you do not retype anything.

This works because you are only paying for the expensive model on the questions that actually needed it, and you found out which those were by trying. It is much better than guessing in advance, and it is far better than the common alternative of paying frontier prices for everything on the grounds that you might.

It has a second benefit. Comparing two models on the same question, back to back, is the fastest way to build an instinct for what each one is like. A month of that is worth more than any benchmark chart.

When to go straight to the top

Escalation is the default, not a law. Go straight to the frontier model when the task has a shape the small models reliably fail: long chains of reasoning where an early mistake poisons everything after it, work that must hold many constraints at once, anything where you genuinely cannot check the answer yourself, and final drafts of things that go out with your name on them.

That is the honest list. It is shorter than most people’s habit.

Pictures and video are a different game

For text the price differences are pennies and the escalation loop is cheap. For pixels the spread is enormous — a picture is hundreds of chat messages and four seconds of video can be thousands — so the same instinct applies, harder. Draft on the cheap model at the low resolution, decide what you actually want, then spend once on the good one.

The mechanics of that, with the real numbers, are in the piece on making pictures and video. The prices for every model are on the pricing page, without an account.

What benchmarks will not tell you

No leaderboard measures the thing you actually care about, which is whether the model is good at your work. Benchmarks are averages over other people’s tasks. The way to find out is to keep one real question of your own — something you know the right answer to and that is hard — and ask it of each model when you are choosing. Ten minutes of that beats an afternoon of reading comparisons.

Then pick per job, not per brand. That is the whole argument.