Per-million-row cost across Haiku 4.5, GPT-5 nano, Gemini 3 Flash, and a self-hosted Llama on a single L4, with prompt caching, batch API, and structured output flags actually turned on. Most teams over-spec the model: 80% of 'AI classification' work is solved by Haiku 4.5 with caching at sub-$0.20 per 1k rows.

Somewhere in your company there is a backlog of a million rows waiting for a label: tickets that need a queue, invoices that need a cost centre, transactions that need a category, scanned documents that need a type. And somewhere there is a meeting where an engineer has a pricing page open in one tab and a spreadsheet in the other.

We have sat in that meeting more times than we would like, and the spreadsheet nearly always asks the wrong question. At a million rows the model matters less than three flags on the API call, and the model most teams pick is a tier or two above the one the job needs. Our position, from the backlogs we have cleared: 80% of what gets called 'AI classification' is solved by Haiku 4.5 with prompt caching on, at sub-$0.20 per 1k rows.

What follows is the comparison we actually run: four candidates, three flags, the self-hosting question and the test that says the cheap tier is enough. Qualitative on purpose: prices move, the ordering does not.

Which model is cheapest for classification at a million rows?

  • Haiku 4.5. Our default, and the one we leave only with evidence. Cheapest of the hosted three once caching is on, and the most conservative on ambiguous rows, which in classification is a feature: an honest 'unsure' costs less than a confident wrong label.
  • GPT-5 nano. Roughly comparable in cost to Haiku with caching, and fine on short inputs with short label sets. A little more willing to commit on thin evidence in our runs, so more rows come back for a second look.
  • Gemini 3 Flash. The same cost band, and the strongest of the three on long inputs. If your rows are multi-page scanned documents rather than short strings, test it against Haiku before you decide.
  • A self-hosted Llama on a single L4. The only candidate with a fixed bill: the GPU costs the same for one row or a million. Cheapest per row only once the volume is steady and somebody owns the box.

The three hosted tiers sit within touching distance; the gap between them is smaller than the gap between caching on and off. Anyone quoting per-row prices to the last decimal is selling a slide. What is not close is the gap to the frontier tiers: Sonnet, Opus and their equivalents cost an order of magnitude more per row, and for a label from a closed set they almost never earn it.

Why do teams over-spec the model?

Three reasons. The demo was built on the largest model because that was the tab the engineer had open, and nobody wanted to make it worse. Classification gets confused with reasoning, so picking one item from a known list is treated as if it needed a model that can draft a legal opinion. And a wrong label feels more expensive than the bill, until the bill arrives.

Invoice coding is the archetype: a supplier invoice lands in an inbox and needs an account and a cost centre, against rules the finance team can already recite. Short input, finite label set, rules that exist. That is the 80%, and ticket routing, transaction categorisation and document type detection are the same shape. The remaining fifth is rows where the label depends on context the row does not contain, or where the label set is still being argued about. No model fixes the second problem; the expensive one produces wrong entries faster.

Most teams over-spec the model and under-spec the labels. The bill comes from the first mistake and the errors come from the second.

Which three flags change the bill more than the model does?

Every candidate above assumes these are on. Most cost horror stories we hear are a good model with all three off.

  • Prompt caching. In classification the prompt is the same for every row: instructions, label definitions, examples. Only the row changes. Cached, you pay full price for that prefix once and a fraction thereafter; uncached, you pay for it a million times. Keep the static part first and the row last, and watch the cache hit rate, because a timestamp or session id in the prefix breaks it silently.
  • The batch API. A backlog is not latency sensitive; nobody is waiting on any single row, so there is no reason to pay the interactive price. Batch endpoints are discounted for exactly that workload and end the rate-limit fight. Teams skip them because the demo was interactive and the production job inherited its shape.
  • Structured output. Ask for a label from a closed set, in a schema, not a paragraph that happens to mention one. Fewer output tokens, no parser, no row where the model explained itself and forgot to answer. And evaluation becomes trivial: the output is in the set or it is not.

Those three, with Haiku 4.5, are what get you to sub-$0.20 per 1k rows. Turn them off and the same model costs multiples of that; turn them on with a frontier model and you are still paying an order of magnitude more than the job needs.

When does a self-hosted model win?

The self-hosted candidate has a different cost shape, not a different price. A hosted API is flat per row; a GPU is a fixed bill, so its per-row cost falls the more you push through it. It wins only when three conditions hold. The volume is steady rather than a one-off clear-out. The data cannot leave your network, for a reason a lawyer will put in writing. And a named person will keep the box patched, the model version pinned and the throughput tuned. Fail any of those and the hosted tier with caching is cheaper the moment you count the engineer. A single L4 is enough hardware; the question is who owns it after the backlog is cleared.

How do you know the cheap tier is good enough?

Measure it, on the ugly rows. Pull a labelled sample from the real backlog, not the tidy examples from the demo, and have the people who do the job today label it. Run the cheap tier and one tier up on the same sample and prompt, and compare both against the human labels. If the cheap tier disagrees with the humans about as often as the humans disagree with each other, it is good enough.

If it is not, look at the disagreements before you switch model. Most are label definitions rather than model failures: two categories that overlap, a label nobody uses any more. Fixing the definitions improves every model on the list. Then route rather than choose: have the cheap tier emit 'unsure' and send those rows to a larger model or a person. That is how the 80% becomes the whole backlog at close to the cheap price. Keep the sample; the cheapest model is the one you can prove has not drifted.

What to do next

Before anyone opens a pricing page, write the label definitions down and pull the sample. Run Haiku 4.5 with all three flags on and read the disagreements. Escalate only when the sample says you must, and only for the rows that need it. Most of the backlogs we are shown never get past that first step, which is the point.

That is what our first call is for: a 20 or 45-minute call, no deck, to hear what is in the backlog and what a wrong label costs you. Where the scope needs defining, a four-day Spec from €5,000 produces a written brief you own outright: the label set, the eval sample, the routing rule and a costed recommendation on model and flags. Vendor-neutral; take it to us or to anyone.

Or skip ahead and talk through it directly