Measuring an AI support assistant: deflection, intent accuracy, resolution time
An AI support assistant is judged on three numbers: what it resolves without an agent, what it understands correctly, and how long a request takes to be settled. Each is simple to state and easy to inflate. Here is how to define and instrument them so they stand up in front of an executive committee.
Why these three metrics, and not others
You can track dozens of metrics on an assistant: conversation volume, satisfaction, click-through rate, cost per query. They are useful for steering. But to decide whether the system creates value, three questions are enough:
- How many requests are settled without human intervention? That is deflection.
- Does the system understand what the person is asking, in the language of the business? That is intent accuracy.
- Are requests settled faster, including those that go through an agent? That is resolution time.
The three hang together. High deflection with low accuracy means you are sending people into a void. Good accuracy with no time saved means the routing achieves nothing downstream.
Deflection: what is actually resolved
Definition
The share of incoming requests handled end to end by the assistant, with no ticket created and no transfer to an agent, and no return from the same person on the same issue within a defined window.
Instrumentation
- A conversation ID linked to the customer ID, or to the property, case or contract.
- An explicit end status: resolved by the assistant, transferred, abandoned.
- A recontact window set in advance, for example seven days, across all channels: chat, phone, email.
The trap
Counting a conversation as deflected because it ended without a ticket, when the person calls back the next day. This is the most common mistake, and it can make a system look excellent when all it does is push the request onto another channel. An abandoned conversation is not a resolved conversation.
Intent accuracy: understanding in the language of the business
Definition
The share of requests for which the intent detected by the system matches the real intent, as established by a human on a reference sample.
Instrumentation
- An intent taxonomy built with frontline teams, not in a workshop among consultants.
- A reference set annotated by people from the business, drawn from real requests and refreshed regularly.
- Measurement per intent, not just overall: the average hides the critical categories.
The trap
Measuring on generic vocabulary. A tenant does not talk like a test set: they use in-house terms, abbreviations and typos, and mix two requests in one message. That is why I always specify “on domain vocabulary”. Accuracy measured on clean sentences is worthless in production.
Second trap: leaving “out of scope” messages out of the calculation. They must be detected as such, and that detection is part of accuracy.
Resolution time: what the user experiences
Definition
The average time between first contact and actual resolution of the request, across all channels and all paths: resolved by the assistant, routed to an agent, or escalated.
Instrumentation
- A timestamp for first contact and for closure, in the same system of record as the baseline period.
- A baseline period before deployment, measured the same way.
- A breakdown by request category, to compare like with like.
The trap
Choosing your scope after the fact. Measuring resolution time only on the requests the assistant resolves on its own, or only on the simple categories, gives a flattering and false number. The right metric covers the whole flow, including the hard cases that go through an agent, and compares comparable periods and request mixes.
An illustration: Home Partners
For the contact centre of Home Partners of America, a residential portfolio of 30,000 units, the AI system qualifies, routes and resolves tenant requests, in production across the whole portfolio.
Engagement performed at Home Partners of America, a Blackstone subsidiary, 2022 – 2023. Public data, informed stakeholders, within applicable confidentiality agreements.
What matters in this reading is the combination. Deflection shows the volume absorbed, accuracy shows that the absorption rests on a real understanding of the requests, and resolution time tells you whether customers get a complete answer faster. These definitions are the ones I recommend; the case illustrates how to read the three numbers together.
Five rules for numbers that hold up
- Fix the definitions before launch, in writing, with the business owner.
- Measure a baseline before deploying. Without a “before”, there is no gain.
- Publish the three numbers together, never just one.
- Have the business do the annotation, and refresh the sample: vocabulary evolves.
- Track the cases that come back: they show where the assistant fails without knowing it.
Where does the organisation fit in?
These metrics are not just a dashboard matter. They force decisions on who annotates, who arbitrates the taxonomy and who corrects course when a category drifts. That is the organisational side, the one that in my view accounts for 70% of a successful AI transformation, and the one people forget when they measure a proof of concept (POC). I go into this in from AI POC to production.
The logic is the same as in product: a metric chosen before building, a baseline, a hypothesis to validate. That is what I describe for market fit, and what should feed the assistant’s improvement backlog. For simpler internal use, ClairAI explains how to measure the time AI saves, counting review and corrections.