Vendor reports compare apples (chat tickets the bot took) to oranges (the email backlog it didn't), and ignore the cost of a degrading escalation queue. We propose a five-metric scoreboard and the audit pattern that catches inflated wins inside 60 days. The honest sustainable deflection rate is 25-40%, not the 70% vendors quote.

Every deflection pilot report we are shown has the same shape: a headline near 70%, a chart climbing from the week the bot went live, and a recommendation to roll out. The heads of support who bring them to us are not fooled by the chart. What worries them is harder to name: the queue their agents work feels heavier than before the pilot, and the report says the opposite.

The something is the denominator. Vendor reports count the chat conversations the bot closed and divide by the chat conversations it saw, leaving out the email backlog it never touched, the phone calls from customers who gave up on the widget, and the escalations that reach a human already annoyed. Apples against oranges, and the cost of a degrading escalation queue nowhere at all.

Our position, after building and auditing enough of these, is that the honest sustainable deflection rate is 25 to 40%, and that is a good result. The gap to the 70% on the slide is the difference between a pilot that removes cost and one that moves it into a queue nobody measures. Below: why the numbers inflate, what the queue costs, the five metrics on our scoreboard, and the audit pattern that catches an inflated win inside 60 days.

Why do deflection pilots overstate their numbers?

Four mechanisms, and most pilots run all four. The first is the channel: the bot is deployed on web chat, the number is measured on web chat, and a customer who bounces off it and writes an email is a deflection on one dashboard and a new ticket on another. Nobody reconciles the two.

The second is the window. Pilots launch into a quiet period, watched closely by the people who bought them, pointed at the categories the bot is best at: order status, password resets, opening hours. The easy tickets go first, and the rate that survives their exhaustion is not the launch-week rate.

The third is the definition. 'Contained' means the customer did not press escalate, not that they got an answer. Someone who closes the tab and rings the next morning is a contained conversation and an unresolved problem; only one of those is on the slide. The fourth is the baseline, reconstructed after the fact from the vendor's tooling, by the vendor.

What does the escalation queue cost you?

This is the cost the report never carries. A bot that takes the simple tickets leaves the humans a harder mix. Handle time rises, not because agents slowed down but because the easy work that flattered the average has gone. Customers who reach a person have already failed once, so they arrive later in their problem and shorter in patience, and first-contact resolution drops with them.

Then the queue compounds. Harder tickets from angrier customers are the ones agents leave over, so attrition rises among the people who can still handle them. Re-contacts from failed bot conversations land as new tickets, so the backlog the bot was meant to clear grows while the dashboard stays green. None of this is visible on the bot's channel, and it is where most of the money went.

A deflection rate that cannot survive its own escalation queue is not a saving. It is a delay with a dashboard.

What is a sustainable deflection rate?

Sustainable means three things. The rate holds after the launch window, once the easy categories are exhausted and the pilot team has stopped watching. It is measured across every inbound channel, not only the bot's. And re-contacts do not rise, so the deflected customers stayed deflected. On that definition, the pilots we have seen hold up land between 25 and 40%.

The 70% figure is not a lie in the narrow sense. It is a containment rate, in the bot's own channel, over the launch window, on the categories the bot was pointed at, and none of those qualifiers appears on the slide. Remove them and the number roughly halves. That is not a failure; it is the real size of the prize, provided the queue behind it is not paying for it.

What are the five metrics on the scoreboard?

One scoreboard, five lines, refreshed weekly from your own ticketing system, not the vendor's dashboard. Each line catches a specific way the headline inflates.

  1. True deflection rate. Conversations the bot resolved with no human contact on any channel inside the re-contact window, over all inbound contacts on all channels. Catches the channel switch.
  2. Re-contact rate. The share of deflected conversations where the same customer came back on any channel inside the window. Catches the deflection that was really a delay.
  3. Escalation queue mix and handle time. What reaches a human, in which categories, and how long it takes. Catches the degrading queue the bot's dashboard cannot see.
  4. Sampled resolution quality. A person grades a weekly sample of bot-closed conversations on one question: did the customer get what they came for? Catches 'contained' that was never 'resolved'.
  5. All-in cost per resolved contact. Bot spend, agent time and rework across the whole operation, over contacts genuinely resolved. Catches the pilot that moved cost instead of removing it.

Vendors resist the first and fifth lines, because both need data they do not hold. That resistance is itself a finding.

How does the audit pattern catch an inflated win inside 60 days?

Sequence is most of the pattern. Freeze the baseline before the bot goes live, from your own system: all channels, all categories, the five metrics computed on it. Then hold back a slice of traffic where the bot is not deployed, a region, a segment or a rota window, and run both sides on the same scoreboard week by week. The held-out slice shows whether a quieter month or a product fix did the deflecting.

Each week, reconcile the vendor's number to the true rate from your own data and grade the sample. Sixty days is the length because the first weeks are a honeymoon of easy categories and close attention. An honest pilot's headline falls towards the true rate over the 60 days and the two lines meet. An inflated one keeps the headline high while the gap widens, re-contacts climb and escalation handle time drifts up. None of this needs a statistician, only five lines on one page and a baseline nobody may edit.

We write the audit into the acceptance criteria before the pilot is priced, and tie the vendor's fee to the true rate and the re-contact rate, not their own containment number. A fee based on a self-reported figure is an incentive to inflate.

What to do next

Build the baseline before you talk to anyone selling deflection. If your ticketing system cannot produce an all-channel count of resolved contacts today, that is the first project, and a smaller one than a bot. Then decide what the bot may own: rules-based questions with a known answer move to the system; judgement, complaints and the customer whose problem fits no category stay with your team, and the scoreboard proves that line holds.

That is the conversation we have on a 20 or 45-minute call, no deck and no homework, ending in a one-page summary of where your numbers stand. Where the scope needs defining, the four-day Spec from €5,000 produces a vendor-neutral brief you own outright: baseline, five metrics, hold-out design and acceptance criteria, before any pilot is priced. We build on the helpdesk you already run rather than selling a bot, so we have no containment number to defend; the average build ships in eight weeks, scoreboard included. A rate between 25 and 40% that holds is worth paying for. One at 70% that exists only on a vendor's dashboard is a cost you have not found yet.

Or skip ahead and talk through it directly