Recall and signal-extraction tested on 200 anonymised B2B sales calls across Sonnet 4.6, GPT-5, and Gemini 3.1. A templated Haiku 4.5 prompt with a structured schema beats default Sonnet output for 80% of summary use cases at one-fifth the cost, but only if the rep can edit the schema.

Ask a room of sales leaders which model their call-summary tool runs on and most will not know. Ask the vendor and the answer is usually the most expensive tier, because that is the one that demos well. Nobody chose it. It was the default, and defaults in AI tooling have a way of becoming line items nobody revisits.

So we tested for a better answer. We took 200 anonymised B2B sales calls, ran them through Sonnet 4.6, GPT-5 and Gemini 3.1 with the plain summarisation prompt most tools ship, and scored the output on recall (did the summary contain what mattered) and signal extraction (did it surface the next step, the objection, the person who can say yes). Then we ran the same calls through Haiku 4.5 with a templated prompt and a structured schema.

The templated Haiku beat the default Sonnet output for 80% of the summary use cases we scored, at one-fifth of the cost. That is the headline. The clause that matters more follows it: the result only holds if the rep can edit the schema. Most tooling being bought right now gets the first half right by accident and the second half wrong by design.

Which model tier is right for call summaries?

For most of what a sales team needs from a summary, the cheap tier, with structure. All three frontier defaults were competent at recall and slipped in the same places: the commitment made in passing, the objection buried inside a question, the next step implied rather than stated. A bigger model guesses at those slightly better. A schema stops the guessing, and the cheap model with a schema found them more reliably than the expensive model without one.

The mistake we see is treating model tier as the quality lever and the prompt as plumbing. In call summarisation it is the other way round: the schema does most of the work, and the tier decides how much you pay for the remainder. Buy the tier for the calls that need it, not for all of them.

Why does a schema beat a smarter model?

A default prompt asks for a summary and gets prose. Prose is written to be read, not to be filed: what the model judged important, in the order it chose. And a sentence that is not there is not something you notice.

A schema turns the same task into extraction against a known list of targets. Next step: who, what, by when. Objections: what was raised, how it was answered. The model no longer decides what your pipeline cares about, because the schema has already told it. Three things follow.

  • Misses become visible. An empty field is a miss you can see, chase and count. A missing sentence in a paragraph is a miss nobody finds until the deal slips.
  • Summaries become comparable. Two hundred prose summaries are two hundred documents; two hundred schema-filled summaries are a table, and a manager can read a table.
  • The output fits the CRM. Each field maps to a property on the deal record, so the summary is written into the pipeline rather than attached to it.
The schema is the product. The model is a line item.

What should the schema contain?

Fewer fields than the first draft. Every field should answer a question a manager asks in the pipeline review or map to a property the CRM already has. If it does neither, it is decoration. We start most engagements from roughly this.

  • What was said. A short narrative, capped in length, in the rep's language rather than the model's.
  • What was agreed. Commitments on each side, stated as commitments, not as themes.
  • Next step. Owner, action and date. If one is missing, the field says so rather than inventing it.
  • Objections and responses. What the prospect pushed back on and what the rep said to it.
  • Buying signals. Budget, timeline, decision process and named people, each marked stated, implied or not discussed.
  • Risks and open questions. What the rep still needs to find out before the next stage.
  • Alternatives mentioned. Competitors, incumbents, or doing nothing.
  • The quote behind each field. The verbatim line the claim came from, so a rep can check it in seconds and a manager can trust it.

The last item is the one most tools leave out, and it earns the rest their credibility. An extracted next step with the sentence it came from is a fact. Without the sentence it is an opinion in a database.

When is the expensive tier worth it?

For the remaining fifth, and those calls are identifiable in advance. The frontier models earned their price when the task stopped being extraction and became interpretation:

  • Multi-party calls where speakers overlap and the commitment that matters was made by the quietest person in the room.
  • Calls where what was agreed is implicit or contradictory, and the summary has to reason about which version the prospect meant.
  • Coaching, where the question is not what was said but how the rep handled it.
  • Mixed-language calls, or calls thick with jargon the schema has no field for.
  • Anything where the summary is the decision, such as a forecast commit read by someone who was not on the call.

Route by call type, not by default. The cheap tier runs everything. The schema raises a flag when it cannot fill a field with confidence, and those calls, and only those, go to the expensive tier. Buying the top tier for every routine follow-up is paying for interpretation you did not ask for.

How do you roll this out without the reps hating it?

Start by letting them own the schema. A schema written by a vendor is right for the average pipeline and wrong for yours: your qualification criteria are not the average, and your enterprise and partner pipelines do not share a definition of next step. A locked schema produces summaries that are correct in general and slightly off in every specific. Reps stop reading within a fortnight, and the tool becomes a record nobody trusts.

Owning it does not mean every rep freelances a personal template. It means one owner per pipeline can add, rename or retire a field, and the change shows up in the next call without a ticket to the vendor. The schema is versioned and reviewed in the weekly pipeline meeting. This is also why the cheap tier holds: when the schema changes, the model does not need to.

The rest is sequencing.

  1. Write the summary into the record within minutes, and let the rep edit it there. One that lands on the deal, with the quote beside each field, is the note the rep would have written anyway.
  2. Start with one pipeline and one owner. The team that argues least about what a next step means goes first.
  3. Read every summary by hand at first, and do not score reps off it until they trust it. Every correction is a schema edit, made the same week. Once the summary feeds a conversation about quota, reps start talking for the model instead of to the prospect.

What to do next

If you are choosing tooling, two questions cut through most vendor decks: which tier runs by default, and whether the schema is yours to edit without a ticket. A vendor that answers the first with the most expensive tier and the second with no is selling you the wrong thing at the wrong price.

If you would rather have this built on the CRM you already run, with a schema your team owns and a model chosen on evidence rather than habit, that is the work we scope. It starts with a 20 or 45-minute call and a four-day Spec from €5,000 that produces the brief, the draft schema and the acceptance criteria before any build is priced. The average build ships in eight weeks and 95% of what we scope reaches production. The judgement on the call stays with your reps. The typing does not.

Or skip ahead and talk through it directly