The Two Directions of AMD Error Are Not Symmetric
AMD accuracy percentages hide which way you err. False-machine drops live prospects silently; false-human burns agent minutes visibly. They are not equal.
The Two Directions of AMD Error Are Not Symmetric
Fifteen to thirty seconds. That is what a false-human error costs when an agent gets bridged into a voicemail: the agent listens, realizes, disconnects, wraps up. It shows up in the handle-time report the same day, and somebody complains about it by lunch. Now guess what a false-machine error costs. A live prospect says hello, the detector says MACHINE, the call drops or routes to a message, and the report says... nothing. The call never reaches an agent. It never enters a conversion funnel. It is a perfectly silent subtraction of revenue. That asymmetry is the whole subject of this piece, and almost every AMD evaluation I have ever seen ignores it.
Everyone reports AMD error as one number. "% accuracy," or its mirror "% error." The moment you collapse two different failure modes into one scalar, you have thrown away the only information that determines what the error costs you. A classifier at 95% with all five points of error in the machine-detected-as-human direction is a different business than a classifier at 95% with all five points the other way, and on a live floor the second one is quietly catastrophic while the first one is merely annoying. Error direction is not a footnote to accuracy. In operational terms, it is the accuracy.
The two errors, defined precisely
Asterisk gives us the clean vocabulary, so let me use it. app_amd sets AMDSTATUS to HUMAN, MACHINE, or NOTSURE, and AMDCAUSE carries the threshold that produced the verdict plus the measured duration in milliseconds (the AMD terms glossary covers the full vocabulary). Two of those verdicts can be wrong about a live answer in opposite ways:
False human on machine (agent hears voicemail)
The call was answered by an answering machine or voicemail system. The detector says HUMAN. The dialer bridges an agent. The agent listens to "Hi, you've reached the Hendersons, please leave a message," waits for it to be over, realizes, hangs up, dispositions the call. Cost: 15 to 30 seconds of agent time per event, plus the morale tax of agents knowing the dialer wastes their attention. Visibility: immediate. It lands in average handle time, in agent talk-time reports, in the queue metrics every supervisor stares at all day.
False machine on human (the silent killer)
The call was answered by a live person. The detector says MACHINE. What happens next depends on your dialplan, and none of the branches are good. The call gets dropped. The call gets a message dropped on a live human, which is a compliance exposure and an insult in one gesture. Or the call sits in the machine path until the human, hearing nothing, hangs up. Cost: the entire value of the conversation you could have had, plus the lead, plus the dial. Visibility: effectively zero in standard dashboards. It never becomes a connect. It is a non-event by construction.
The default Asterisk AMD behavior on real traffic drops an estimated 10-20% of live humans as machines. That is our measured claim from detection data, not a figure from the Asterisk project, and I want to be precise about its provenance because it is exactly the kind of number that gets laundered into folklore. Whether your installation is at the low end or the high end depends on your carriers, your amd.conf tuning, and your traffic mix, which is why the measurement procedure at the end of this piece matters more than my estimate.
Why the default tuning errs toward the machine side
The amd.conf defaults explain the skew mechanically. maximum_number_of_words is 3: a human who answers with "Hi, this is Maria from the doctor's office, how can I help you" has already exceeded the word budget and forced the detector into its fallback thresholds. greeting is capped at 1500 ms, which a long professional greeting outruns. initial_silence of 2500 ms means a hesitant answerer who pauses before speaking is treated to two and a half seconds of the detector leaning MACHINE. Every one of those defaults is a reasonable individual choice, and their composition points the same direction: calls that deviate from a terse "hello" drift toward MACHINE. The deviation profile of real humans, chatty greetings, pauses, background noise, is exactly the profile the defaults punish.
Why one error screams and the other whispers
The economics of the two errors are not just different in size. They are different in sign, in who notices them, and in when they get noticed. That difference explains a decade of misallocated AMD tuning effort across this industry.
The visible error gets fixed; the invisible error gets ignored
A supervisor sees agents stuck on voicemails and opens a ticket the same morning. The fix everyone reaches for is making the detector more aggressive about declaring MACHINE: raise the thresholds, wait for more certainty before committing to HUMAN, let after_greeting_silence run longer. Every one of those changes reduces false-human errors, and every one of them increases false-machine errors, because a threshold detector moving its decision boundary trades one direction for the other. The room feels better. Handle time drops. Meanwhile the connect rate quietly sags, and nobody connects the sag to the tuning change, because the calls the tuning killed never generated a record anyone reads.
This is the trap: the error that produces complaints gets tuned away, and the payment for tuning it away is extracted from the error that produces no complaints. I have watched rooms do this to themselves for months at a time. The dashboard says AMD accuracy improved, handle time improved, and the campaign's revenue per hour decays for no reason anyone can name. The reason has a name. It is in the confusion matrix nobody prints.
Connect-rate dashboards hide the drop, until you audit
Here is the mechanism by which false-machine errors stay invisible, and it is worth slowing down for. Your connect rate is live connects divided by answered calls, or divided by dials, depending on the dashboard. A false-machine event lowers the numerator, but it also, in most reporting, removes the call from the "live opportunity" population entirely, because the reporting layer believes the detector: if AMDSTATUS said MACHINE, the call was a machine. The error corrupts its own denominator. The dashboard is not failing to show the damage; the dashboard is actively agreeing with the thing that caused it.
So the drop in performance attributable to false-machine errors surfaces as a diffuse malaise: lower contact rate, lower conversion per dial, "the data is worse this month." Attribution requires an audit, and the audit requires listening to calls the detector labeled MACHINE. Which almost nobody does as a routine matter, because there are thousands of them and the ones that matter sound like this: ring, ring, "hello? ... hello?" click. Four seconds of audio. The most expensive four seconds on your dialer, filed under machine.
The cost table, without invented numbers
I am not going to hand you a spreadsheet of benchmark costs, because a cost-per-error number is only real when computed from your own conversion economics. What I can give you is the honest structure of the calculation, and where each quantity comes from:
| False human on machine | False machine on human | |
|---|---|---|
| What happened | Voicemail bridged to an agent | Live prospect dropped or dumped to message |
| Immediate cost | 15-30 s agent time, ~1 event per misdetection | Full value of the lost conversation (your revenue per live connect) |
| Where it shows up | Handle time, agent talk time, same-day complaints | Nowhere standard; looks like lower contact/conversion rate |
| Who notices | Agents and supervisors, immediately | Nobody, until a recording audit |
| Typical tuning response | Push thresholds toward MACHINE (which worsens the other direction) | Usually unaddressed; often not even measured |
| How to measure it | Review a sample of HUMAN-verdict recordings, count actual machines | Review a sample of MACHINE-verdict recordings, count live humans ("hello?" with no beep) |
| Compliance angle | Wasted agent time, no regulatory exposure | A message dropped on a live human can count toward TSR abandonment problems |
One correction to that middle row on measurement, because precision matters here: the false-human rate is measured by reviewing calls your detector labeled HUMAN and counting actual machines. The false-machine rate is measured by reviewing calls labeled MACHINE and counting actual humans. Both directions require human review of a sample, or a second classifier used as reference with its own error acknowledged. There is no dashboard shortcut, and anyone selling you a shortcut is selling you the same blind spot this article is about.
The revenue side of the false-machine calculation is yours to fill in: live connects per day, revenue per conversion, conversions per live connect. When operators actually run this arithmetic on a 10% false-machine rate, the number that comes out is routinely larger than their entire dialer infrastructure budget, which is the economic fact the whole AMD-category debate rests on. Do not take my word for the magnitude. Take the afternoon and compute yours.
Where NOTSURE hides the skew
The third verdict deserves its own section, because it is where asymmetric errors go to launder themselves. When app_amd exhausts its analysis window without a threshold being crossed cleanly, AMDSTATUS comes back NOTSURE and the routing decision falls to your dialplan. We have written elsewhere about why the NOTSURE bucket skews toward live humans; the mechanism is simple. Answering machines are statistically easy: long greeting, beep, silence. Live humans are hard: "Hello?" ... "Hello?" ... short words, long uncertain pauses, background noise. The calls that defeat the thresholds are disproportionately the quiet, hesitant, real people, which means your NOTSURE bucket is a live-prospect bucket wearing a machine costume.
What your dialplan does with NOTSURE decides which error direction those prospects become. Route NOTSURE as machine and you have chosen to drop your most-difficult live humans, concentrating false-machine error on exactly the population that was hardest won. Route NOTSURE as human and you have chosen to feed agents the residual voicemails, inflating false-human error. There is no neutral setting. There is only a choice you should make deliberately, with numbers, rather than by inheriting whatever the dialplan template shipped with. The share of calls landing in NOTSURE is itself a metric worth trending; a detector whose verdict begins at 125 ms at 99% accuracy, which is the spec we built AMDY around (see the features page), spends far less time parked in undecided territory than a threshold detector waiting out a 5000 ms total_analysis_time window with a 2500 ms initial_silence and an 800 ms after_greeting_silence in front of it.
The confound that corrupts both directions: FAS
Before you run the audit, one contamination source needs a name. False answer supervision, FAS, is a carrier signalling a call as answered when no party picked up: a ring, a billing timer, and silence where a human or a greeting should be. Every downstream consumer of that signal inherits the lie. Your detector analyzes audio that contains no answer at all, and whatever verdict it produces is wrong by construction, in a direction that has nothing to do with the classifier's actual behavior. FAS does not merely add noise to your error-rate measurement. It adds biased noise, because a silent FAS call looks acoustically like nothing, and "nothing" gets classified toward MACHINE by most threshold logic, inflating your apparent false-machine rate with events that were never live humans in the first place.
The practical consequence for the audit in the next section: your recording reviewers need a fourth truth label. Not just live person, voicemail, ambiguous, but also "no answer event present," for calls where the audio is dead air or ring-through from end to end. If your detection log shows a verdict on a call with no audible answer, that is a carrier problem wearing a detection costume, and it belongs in a per-carrier report, not in your classifier's confusion matrix. We surface a per-carrier FAS breakdown in the product for exactly this reason: until you separate carrier-side false answers from classifier-side errors, neither number means anything. We cover the mechanism in more depth in our piece on carrier false answers, and if the audit below shows your false-machine rate varying wildly by carrier trunk, FAS is the first suspect, not your amd.conf.
The regulatory echo
Error direction is not purely an economics question. Under the FTC Telemarketing Sales Rule, abandonment is capped at 3% of live answers per campaign over a rolling 30 days, per 16 CFR 310.4(b)(4). A false-machine verdict on a live human, followed by a dropped message or dead air, behaves in that accounting like an abandoned call that you orchestrated on purpose. A false-human verdict on a voicemail has no such reading; it wastes your money and nobody else's. So the two error directions differ not only in visibility and cost but in regulatory character: one is an internal inefficiency, the other brushes against an external enforcement metric. If your AMD tuning strategy is "when in doubt, call it a machine," you have tuned your compliance posture along with your handle time, and possibly in the wrong direction on both. The safe-harbor rule is the FTC's own document and worth reading directly rather than trusting my paraphrase of it.
How to measure your own error directions
The procedure is an afternoon of work and it is the only AMD number I would trust from a vendor or from myself. Our full accuracy-testing procedure covers the general method; here is the error-direction core of it.
Step one: pull a stratified sample of recordings
From one week of traffic, sample three buckets: calls labeled HUMAN, calls labeled MACHINE, and calls labeled NOTSURE. Random within each bucket, a few hundred each if your volume supports it. The stratification matters, because the population shares differ wildly and a flat random sample would undercount whichever verdict is rare.
Step two: label them by ear
A human listens to the first ten seconds of each recording and tags the truth: live person, voicemail greeting, or genuinely ambiguous. Yes, this is manual. Yes, it is tedious. The ambiguity in AMD evaluation comes from the data, and pretending the labeling step can be skipped is how vendors end up quoting accuracy numbers nobody can reproduce.
Step three: build the two-by-two and read the directions separately
Cross tabulate verdict against truth. You get four cells: correct-human, correct-machine, false-human-on-machine, false-machine-on-human. Now report the two error rates separately, as percentages of their own truth populations: of all actual humans, what share did we call machines; of all actual machines, what share did we call humans. Two numbers. Track them over time and after every tuning change, because as established above, most threshold tuning silently trades one against the other, and the trade shows up in this table before it shows up anywhere else.
Step four: monetize both directions with your own economics
Multiply the false-machine rate by live answers per day and revenue per live connect. Multiply the false-human rate by machine answers per day and the fully loaded cost of 15-30 seconds of agent time. Compare the two products. In my experience the false-machine product dominates, usually by an order of magnitude (we traced the full cost chain in the real cost of AMD false positives), but your mix may differ, and the point of the exercise is to replace my experience with your arithmetic. This is also the calculation to redo before buying any detector: ask the vendor for their error rates by direction, and note carefully whether they even understand the question. A vendor who answers "we're 97% accurate" to a question about error direction is telling you they have never thought about which of your dollars they are spending.
The takeaway
If your AMD reporting has one accuracy number, it is hiding the only distinction that determines what errors cost you. Split it. The error that generates agent complaints is priced in agent seconds and gets fixed by lunch; the error that generates silence is priced in entire conversations and gets fixed never, unless you go looking for it with your own ears on your own recordings. Split the number, price both directions, and tune deliberately.
The question I would genuinely like answered by anyone running this audit on their own floor: what did your false-machine rate turn out to be, and how long had the dashboard been assuring you everything was fine?