← All articles
Product UpdatesOct 9, 2026 13 min read

The KPI Gap Between 'Answered' and 'Actually Spoke to a Human'

Most call centers report answered calls, not real human conversations. Learn to compute delivered-live-conversation rate from data you already own.

The KPI Gap Between 'Answered' and 'Actually Spoke to a Human'

There is a number on your dashboard called "answered calls," and there is the number that actually pays your agents: conversations with a live human. Those two numbers are not the same, and in most centers nobody computes the second one. That is the KPI gap, and it hides your worst leak, your riskiest dialer behavior, and your best argument for budget, all in one place.

You can compute it Monday, from data you already have. No new hardware, no vendor, no project plan. I will show you the exact arithmetic, what a healthy gap looks like, and the three decisions the number should be driving.

The two numbers every center conflates

Open any outbound dashboard and you will find some version of these metrics:

  • Dials attempted. Every call placed.
  • Answered. The far side picked up. A person, a voicemail box, or something in between.
  • Handled or connected. The call reached an agent queue or message was played.

Answered is where the trouble starts, because "answered" merges two completely different events. A voicemail greeting answering is an "answer." A live human saying hello is an "answer." A carrier intercept tone is often an "answer." The dashboard counts them identically, and then everyone downstream, the agents, the managers, the board, works from a number that quietly overstates reality.

The number that matters I call delivered-live-conversation rate: the share of answered calls where a live human was correctly identified, routed to an agent, and actually spoke. It is the point of the entire operation. Everything before it is cost.

Here is the uncomfortable part. The default answering machine detection in most dialers drops an estimated 10 to 20 percent of live humans, measured on real traffic. Those are people who said hello and got a hangup or dead air. Your "answered" count includes every one of them. Your agents never saw them. Your revenue never saw them. Your compliance exposure did, because a human hung up on by a machine is, in the eyes of the FTC's telemarketing rule, an abandoned call counting against the 3 percent abandonment cap per campaign per 30 days. The rule itself is public at ecfr.gov, and the FTC's enforcement history sits at ftc.gov.

So the gap between answered and spoke-to-human is not a rounding artifact. It is misclassified humans, silent voicemail routing, and slow verdicts, all pooled together and reported as one happy number.

Where the gap physically comes from

Three leaks, in order of size.

Leak one: humans classified as machines

What the leak costs, in people

The dialer's AMD listens for the first fraction of a second and makes a call: human or machine. When it guesses machine on a human, the call is dropped or parked on a recorded message. The person hears a click. Your abandonment count ticks up. Your answered count already ticked up. Nobody talks to anyone.

A 10 to 20 percent human-drop rate sounds abstract until you put it on real volume. On 100,000 answered calls with a 30 percent live-human rate, you have 30,000 humans. At 10 percent misclassification, 3,000 of them got the hangup this month. At 20 percent, 6,000. That is the gap, expressed in people.

Which direction your errors run matters as much as the rate. Machines classified as humans waste agent seconds and inflate your connect numbers. Humans classified as machines burn compliance budget and goodwill. Most default settings err in both directions at once. We break the tradeoff down in AMD error directions and which one you can afford.

Leak two: slow verdicts that eat the greeting

Detection takes time, and every millisecond of it is dead air on the line. A verdict that starts around 125 milliseconds, which is where a fast detection engine lands, is imperceptible. A slow one produces that half-second of silence every consumer has learned to read as "this is a robocall," and people hang up before an agent can pick up. The call shows as answered. No conversation occurs. Latency versus accuracy is a real engineering tradeoff and we walk through it in AMD latency versus accuracy.

Leak three: drift

The classification threshold that was right in January is wrong by July. Voicemail greetings change, new handsets enter the pool, carrier behavior shifts. The gap widens slowly enough that no one flags it, and the answered metric keeps looking fine the whole way down. This is model drift, and it is the reason a one-time tuning is not a strategy: AMD model drift.

How to compute delivered-live-conversation rate, Monday morning

You need three counts out of your existing systems, per campaign, for a fixed window like last 30 days:

  1. Answered calls, from your dialer's call log. The raw event count, nothing adjusted.
  2. Calls where an agent spoke with a live human. In most dialers this is the disposition count minus machine dispositions minus dropped-in-queue. If your agents disposition honestly, this number already exists.
  3. Calls your AMD classified as human but the agent found dead or voicemail. This is the false-positive count. Some dialers expose it; if yours does not, agent dispositions of "no one there" against an AMD human verdict approximate it.

Then:

Delivered-live-conversation rate = conversations with a live human / answered calls.

Run it per campaign, not company-wide. The FTC abandonment cap is per campaign per 30 days, so per-campaign is also the lens compliance requires. While you are there, compute the companion number: humans dropped by AMD divided by total live humans detected plus dropped, which is your effective human-drop rate. If you cannot get it from your dialer logs, a month of per-call detection logging will hand it to you directly, and that same log is what you show a regulator: how AMD logs become your compliance defense.

Total effort: a report pull, a join, two divisions. If it takes longer than an hour, your reporting stack owes you an explanation.

The KPI scorecard

Here is the scorecard I would put in front of a leadership team. Rates and structure only, no invented benchmarks; fill the cells with your own pull and let the campaign rows fight it out.

Metric Definition What it tells you Where the data lives
Answered rate Answered calls / dials List and timing quality Dialer call log
Delivered-live-conversation rate Human conversations / answered calls The real funnel, post-answer Dialer log + agent dispositions
Human-drop rate Humans misclassified as machine / all live humans Your AMD's cost in lost conversations AMD verdict log vs. dispositions
False-machine-to-agent rate Machines sent to agents / agent-handled calls Agent time wasted on voicemail Agent dispositions
Abandonment headroom Abandoned calls vs. the 3% per-campaign/30-day cap Regulatory runway Dialer log + the FTC rule
Verdict latency Time to machine/human verdict Dead air your callers hear AMD engine reporting

Six rows, one page, and a completely different conversation at the next ops review.

The Monday pull, step by step

Since I keep saying this takes an hour, here is the hour.

Step one, pick the window. Last full 30 days, per campaign. Not quarter to date. Thirty days matches the abandonment measurement period, which means the compliance number and the ops number come from the same denominator and nobody argues about mismatched windows later.

Step two, export three raw counts per campaign. Dials, answered, and every agent disposition with its call id. Do not let anyone pre-aggregate for you. Pre-aggregated reports are where misclassified humans go to disappear, because the aggregation usually keys off the AMD verdict, which is the thing you are auditing.

Step three, join dispositions to call records. Most dialers already store this link. The join gives you, per call, what the AMD said and what the agent found. Every disagreement is one of your two error directions, and you can count them without trusting either side alone.

Step four, compute the two rates. Delivered-live-conversation rate and human-drop rate, per campaign. Put them side by side with answered rate. Three columns, one row per campaign.

Step five, sit with the outliers. The interesting campaigns are not the worst ones, they are the ones where answered rate looks great and conversation rate does not follow. That divergence is a detection problem masquerading as a list problem, and it is almost always the campaign where someone recently changed a dialer setting or imported a new list source without telling you.

Then stop. Do not fix anything on day one. The first week of the number is for establishing that it is stable and that people believe it. Fixes start week two, and they start with the campaign that has the most conversations to recover, not the loudest manager.

A note on whose numbers you are trusting

The whole exercise above assumes your dialer's own verdicts are available for comparison, and they are, but availability is not accuracy. If the AMD says human on a call where no human existed, your agent disposition catches it. If the AMD says machine on a human, your log shows nothing at all, because no agent ever saw the call. The second error is invisible in your own data, which is precisely why the estimated 10 to 20 percent figure surprises operators who check their internal stats and find nothing wrong. Nothing is wrong with the stats. The stats never saw the people.

That asymmetry is also why list hygiene alone cannot close the gap. You can scrub a list to gold and still hang up on a fifth of the humans who answer it, because the loss happens after the answer, in the half-second your detection spends deciding. External detection with per-call logging is the only way to count what your own dialer structurally cannot. And before you dismiss the risk of that blind spot entirely, know that some of the numbers on your lists are not consumers at all, they are planted lines run by people who sue for a living, and they love a dialer that cannot tell a human from a machine. More on that in honeypot detection.

What a healthy gap looks like

I am not going to hand you a universal target percentage, because anyone who does is guessing at your traffic mix. What I will give you is the shape of a healthy gap versus a sick one.

A healthy gap is small, stable, and explainable. Answered minus conversations tracks mostly to voicemail volume, which tracks to your list's demographics and the time of day. It moves predictably. When it jumps, you can name the cause: a new list source, a changed greeting pattern, a dialer software update.

A sick gap is wide, drifting, and mysterious. Nobody can say why conversations lag answers by that much. It grew over two quarters without a decision being made. The team has developed folk theories ("summer," "the new list") that survive because no one has the per-call data to kill them. That is a shop being managed off a number that lies, and the fix is not effort, it is measurement.

One more distinction worth the ink: a wide gap hurts in two currencies at once. Lost conversations are lost revenue, and dropped humans are compliance incidents. The same misclassification event debits your top line and your regulatory headroom simultaneously. If you want the full financial framing, including how it plays into cost per number as a life metric, we built a walkthrough on the ROI page.

The decisions the number should drive

Computing a KPI is theater unless something changes when it moves. Three decisions belong to this number.

Decision one: list spend. You buy numbers or data at some cost per record. If delivered-live-conversation rate is 55 percent on one list source and 70 percent on another, identical answered rates are irrelevant. You are paying for answers, not conversations. The gap turns list buying from a volume decision into a yield decision.

Decision two: staffing and pacing. Agent capacity should be planned against predicted conversations, not predicted answers. Shops that staff to answered-call forecasts systematically overstaff, then mask it with pacing that creates queue abandonments, which spends the same 3 percent regulatory budget your AMD errors are already eating. Two leaks, one cap.

Decision three: AMD spend. Once you see the human-drop rate as a line item in lost conversations, detection stops being a dialer setting and becomes a budget question. Modern detection as a service is cheap relative to what it recovers: a free Sandbox tier of 50,000 detections a month with a hard cap and no card to start, a $79 Starter that includes 500,000 detections a month, $299 Growth with 5 million included, and $999 Scale with 25 million, with overage rates that fall per detection at each tier and unlimited servers throughout. Installation is one command by whoever administers your dialer, about five minutes of their time. Full detail is on the features page and the pricing page.

Run the arithmetic on your own gap before you believe me. Multiply your monthly answered calls by your live-human share by the human-drop rate you measure. That product is people who wanted to talk to you. Price the fix against it. If the fix loses on your numbers, you measured wrong.

Where the gap bites hardest

Not every outbound shop loses the same amount to this gap. It bites hardest where three conditions line up: long average handle time, high revenue per conversation, and heavy regulatory scrutiny. That combination describes insurance outbound almost exactly, which is why insurance shops were among the first to treat delivered-live-conversation rate as a board-level metric rather than a dialer curiosity. When one recovered conversation is worth hundreds of dollars and one dropped human is both a lost sale and a potential complaint, the arithmetic of ignoring the gap collapses fast. The insurance case is worth reading even if you sell something else, because it shows the full playbook under maximum pressure: AMD for insurance outbound calling.

If your vertical has shorter calls and thinner margins, the gap still costs you, just in a different currency. High-velocity shops feel it as agent utilization: every machine routed to an agent and every human dropped to a hangup is a seat doing nothing billable, which means the gap quietly sets your effective cost per conversation. Either way, the number is the same number, and you are already generating the raw material for it every day.

One caution before the Monday pull. When the rate first goes up on the board, expect resistance from whoever owns the answered metric, because the new number makes the old one look inflated, and dashboards are territory. The clean defense is that answered still gets reported, unedited, right next to it. Delivered-live-conversation rate does not replace answered any more than profit replaces revenue. It just tells you what the revenue actually was.

Start the clock

The gap between answered and spoke-to-human is the one number in an outbound center that touches revenue, compliance, and agent economics at the same time. It is computable in an hour from logs you already own, and the act of computing it usually settles arguments that have run for quarters.

Pull last month. Divide. Put the rate on the board next to answered, and watch which one people start quoting.

The 50,000-detection Sandbox month, no card, at amdy.io/auth/signup, gives you the per-call verdicts to make the second division honest. You already have the first number. How big is your gap?