← All articles
Product UpdatesSep 28, 2026 13 min read

In Production, Latency Is Accuracy

A 99% accurate AMD model that needs four seconds loses to a 96% model that answers in one. Why time-to-decision governs dialer performance.

In Production, Latency Is Accuracy

Five thousand milliseconds. That is the default total_analysis_time in Asterisk's amd.conf. Five full seconds of analysis window before app_amd gives up and hands your dialplan a verdict. Ask most dialer operators what their AMD accuracy is and they will quote you a percentage. Ask them what their time-to-decision is and you get silence, because almost nobody measures it. That is backwards. In a production outbound room, latency is accuracy.

Here is the uncomfortable version of the claim: a 99%-accurate detector that needs four seconds to decide loses, on the numbers that pay your agents, to a 96%-accurate detector that decides in one second. Not because 96 beats 99. Because the four seconds of dead air, pacing distortion, and abandonment pressure do more damage than three points of raw classification accuracy ever will. Everyone treats AMD as a classification problem. It isn't. It is a classification problem wrapped inside a timing problem, and the timing problem is the one that shows up on the floor.

Why the accuracy benchmark is the wrong benchmark

Walk through what a latency-blind accuracy claim actually measures. Someone takes a corpus of recorded calls, human answers and machine answers, runs the classifier offline, and counts how often the label matches ground truth. Fine. That is a perfectly good offline metric and a perfectly bad production forecast, for one reason: it removes the clock.

In production, the classification does not happen in a vacuum. It happens while a live human is on the line, saying "hello?" for the second time, wondering why nobody is talking. Every millisecond the detector spends deciding is a millisecond the prospect spends in dead air. The accuracy percentage tells you how often the verdict is right. It says nothing about what the call looks like while the verdict is being computed, or what the verdict's lateness does to the call's outcome.

There is a second problem, subtler. Slow detectors inflate their own accuracy numbers. If you allow yourself five seconds of audio, classifying "human vs machine" genuinely gets easier, because by four seconds in, a voicemail has almost always announced itself ("leave a message after the beep") and a human has almost always fallen into frustrated silence or hung up. The accuracy you gain from a long analysis window is accuracy purchased with the prospect's patience. When a vendor quotes a high accuracy figure, your first question should be: measured over how much audio, and at what decision time? If they cannot answer both halves of that question, the number is marketing, not measurement.

A detector that begins returning its verdict at one-eighth of a second, 125 ms, at 99% accuracy, is playing a different sport from one that wants half the amd.conf default window to warm up. We built AMDY's classifier to start deciding at 125 ms precisely because the offline accuracy race was the wrong race. The right spec is accuracy at a decision time, stated together, or it is not a spec.

What slow AMD actually costs on a dialer floor

Let me trace the damage paths one at a time. None of them show up in an accuracy benchmark. All of them show up in your per-hour numbers.

Dead air kills the conversion before the classifier ever runs

The best-studied number in outbound is the two-second rule. Callers who experience silence longer than about two seconds after answering are substantially more likely to abandon the call immediately, and every regulatory regime that touches predictive dialing is built around that human impatience. The FTC Telemarketing Sales Rule, 16 CFR 310.4(b)(4), caps abandonment at 3% of live answers per campaign over 30 days. Abandonment, in that rule, means a live human answers and no agent is there within two seconds. That is the regulator's own estimate of how fast a human bails on dead air.

Now map your AMD onto that. If your detector needs three or four seconds of audio before it commits, every live answer sits in silence for three or four seconds. Not occasionally. Every one. The human says hello, hears nothing, says hello again, and hangs up. That hang-up is an abandoned call in your TSR math. Slow AMD manufactures abandonment out of calls that were already answered, which is the most expensive possible way to burn a dial. You paid for the lead, you paid for the ring, the prospect picked up, and then your detector talked itself out of the conversation while the prospect listened to nothing.

Dialer pacing compounds the latency

Predictive dialers pace on the gap between connect events and agent availability. Slow AMD inserts a multi-second fog into that loop. The dialer cannot count a connect until it knows whether the connect is a person or a machine, so a three-second decision time is a three-second lag in the pacing signal. The dialer's answer to a lagging signal is the same as any control system's answer to stale feedback: it overshoots. It dials ahead to keep agents busy, machines get dropped into the queue or agents get bridged into voicemails, and the room's rhythm degrades in ways agents describe as "the dialer feels weird today."

The signal the dialer actually sees

Worth being precise about what the pacing engine consumes. It does not see your AMDSTATUS string directly; it sees connect events, and in most ViciDial configurations the classification determines whether a connect event counts as a live conversation at all. So decision latency does not merely delay the pacing signal, it gatekeeps which calls exist as pacing input. A slow classifier is not a low-pass filter on your dialer's feedback loop. It is a gate that opens late and sometimes never, which is why pacing symptoms from latency show up as overshoot and queue weirdness rather than as a clean uniform slowdown.

None of this is visible in the accuracy percentage. It is visible in average speed to answer, in agent handle time, in the ratio of live connects to voicemail bridges. Those are the metrics that actually move when you change detection latency, and they are the ones to watch. We keep a standing list of them, and if you take one thing from this piece, take the instrumentation habit: log decision time per call, not just verdict.

Voicemail drop timing collapses

The machine side of the ledger suffers too. A slow detector that waits for certainty before declaring MACHINE has to hold the call open, and when it finally commits, the beep may already be two seconds in the past. Your dropped voicemail message starts mid-word or, worse, the beep passes entirely while the detector was still deliberating. Ask anyone who has run a message-drop campaign on default AMD settings: the percentage of drops that land cleanly is a latency statistic wearing an accuracy costume.

The Asterisk defaults, read as a latency document

The cleanest way to see how the industry thinks about AMD timing is to read Asterisk's own defaults. These are documented values from the Asterisk amd.conf defaults, not my measurements, and they describe a detector designed for patience, not speed:

amd.conf parameter Default What it controls Latency consequence at defaults
total_analysis_time 5000 ms Maximum time app_amd will analyze before giving up Worst-case verdict delay of five seconds on any call
initial_silence 2500 ms Silence before the greeting that triggers an immediate MACHINE A quiet human answering can sit 2.5 s before any verdict logic engages
greeting 1500 ms Maximum length of a greeting treated as human Humans with long greetings ("Hi, thanks for calling Acme, how can I help") risk misreads
after_greeting_silence 800 ms Silence after the greeting before declaring MACHINE Adds 0.8 s to every post-greeting decision
min_word_length 100 ms Shortest burst counted as a word Filters clicks, no latency cost by itself
between_words_silence 50 ms Pause that separates counted words Grows the word-count window on chatty humans
maximum_number_of_words 3 Words before declaring HUMAN Humans who say more than three words force fallback to other thresholds
silence_threshold 256 Amplitude threshold for "silence" Mis-set values silently corrupt every timing above

Read the first row again. Five thousand milliseconds is the default ceiling for a decision. Under the FTC's two-second abandonment lens, a detector that may legally take five seconds to decide is a compliance instrument that cannot help but generate abandonment on its own live answers. And when app_amd cannot resolve the call inside that window, it does not pick the safer branch. It sets AMDSTATUS to NOTSURE and hands the routing decision back to your dialplan, which is where the 10-20% of live humans that default AMD drops (that figure is our measured claim, not an Asterisk admission) leak out of the funnel. We wrote a whole separate piece on why the NOTSURE bucket skews toward live humans; the short version is that machines are easy to classify and quiet, hesitant humans are not, so the undecided pile is full of exactly the prospects you paid to reach.

For completeness: AMDSTATUS ends up HUMAN, MACHINE, or NOTSURE, and AMDCAUSE carries the threshold that produced the verdict plus the measured duration in milliseconds. If you want to know how long your own Asterisk AMD actually takes to decide, that measured-duration field in AMDCAUSE is sitting in your logs right now. Pull it. Distribution, not average. The average is meaningless when the tail is what generates abandonment.

Why a 96% fast model beats a 99% slow one, on the math

Here is the arithmetic that made me stop caring about raw accuracy rankings. Run the numbers on a hypothetical room: 10,000 live human answers per day, agents converting at whatever your rate is, and a classifier at 99% accuracy with a four-second decision time.

At four seconds of dead air, you are not getting 99% of the value of those connects. Some large share of the humans hang up inside the silence window before the verdict ever lands. Call it the dead-air abandon rate. Even if the classifier is perfectly right about every one of them, right about a call the prospect just left is not a conversion. So the effective yield is 99% times (1 minus dead-air abandon rate), and the dead-air abandon rate at four seconds is not small. That is before the pacing distortion, before the voicemail-drop timing damage, before the TSR abandonment the dead air itself creates.

Now the 96% model deciding in one second. One second is inside the human patience window. Dead-air abandonment drops toward zero. The pacing loop gets a clean, fast signal. Voicemail drops land on the beep. You eat a 4% error rate on classification, split across the two error directions, and the dominant costs in the system, the ones that scale with every single live answer, mostly vanish. The 96/1-second configuration wins unless your dead-air abandon rate at four seconds is under roughly three percent, and no room I have run or audited has ever been under three percent at four seconds of silence.

The honest caveat cuts both ways: latency and accuracy are not actually independent knobs. Speed is bought with audio, and audio buys certainty. The whole engineering game, the reason AMDY exists as a server-side classifier rather than another amd.conf tuning guide, is that a model can learn to decide fast on less audio than a threshold algorithm needs. But when you evaluate any detector, yours or a vendor's, you must evaluate the pair. Accuracy at a decision time. One number without the other is half a spec.

How to benchmark time-to-decision on your own traffic

You do not need to trust anyone's marketing numbers, mine included. Measure your own. We publish a full procedure for testing AMD accuracy on your own traffic; here is the latency half of the same discipline.

Step one: extract decision times from what you already log

If you are on Asterisk or ViciDial, AMDCAUSE's measured duration is per-call decision data you already own. Export a week of it. If you are evaluating a replacement detector, the same log export gives you its decision-time distribution side by side. Plot percentiles: median, p95, p99. The median will flatter every detector. The p95 is where abandonment lives, because abandonment is a tail event produced by the slow calls, not the average one.

Step two: split decision time by true label

Decision time is not one distribution. It is two: how fast the detector says MACHINE on actual machines, and how fast it says HUMAN on actual humans. A detector can have a great average by deciding machines fast and humans slowly, which is exactly the wrong trade, since the human side is where dead air costs you the call. Tag every decision time with the verified outcome from call recording review. This is tedious for one afternoon and it is the only method that produces a number you can act on.

Step three: correlate with abandonment and speed-to-answer

Now cross the decision-time data against your abandonment rate and average speed to answer, day over day or, better, campaign over campaign. If you cannot show me a curve where slower p95 decision time associates with higher abandonment in your own data, then your room has different dynamics than every room I have sat in, and you should trust your data over my experience. That is the point of measuring. The dashboard numbers that matter here are the ones we enumerate in our outbound AMD metrics piece, and decision-time percentiles belong on that list even though almost nobody puts them there.

Step four: re-run after any config change

Every amd.conf tuning session trades latency for accuracy, whether or not the operator notices. Raise initial_silence to catch quiet humans, and you have added that raise to every quiet answer's verdict time. Anyone who has tuned these thresholds knows the seesaw. The discipline is to re-measure both sides of the spec after every change, because a tuning that improves the confusion matrix while silently adding a second of decision time is usually a net loss in production, and your accuracy-only dashboard will report it as a win.

What fast looks like when it is real

Concretely, so this is not abstract sermonizing: AMDY's verdict begins at 125 ms, one-eighth of a second, at 99% accuracy. That is the paired spec. Not 99% eventually, not 99% given five seconds of audio. Ninety-nine percent with the clock starting at 125 ms. Install is one curl command on an existing ViciDial box (the wiring is documented here), about five minutes, extension 8370, no dialer-script changes, and the Sandbox tier is free at 50,000 detections a month with a hard cap and no card. You can measure our p95 against your current setup on your own traffic in an afternoon, which is the only evaluation I would accept if I were the buyer, and the only one we ask for.

The pricing scales with volume, not accuracy-at-a-price: Starter is $79/month with 500K detections included and $0.00025 per detection over, Growth is $299/month with 5M included at $0.00015 over, Scale is $999/month with 25M included at $0.00010 over. Unlimited servers on every plan, because charging per server on a latency product would be a strange way to reward people for instrumenting their whole floor.

The takeaway

Accuracy percentages are offline facts about classifiers. Rooms are run on timing. When someone quotes you an AMD accuracy number, the only correct follow-up is "measured at what decision time, on how much audio, and what does your p95 look like split by true label." If the answer is a smaller accuracy number, you have found someone honest. If the answer is a pause, you have found the reason your dialer feels weird.

The one-line version, worth keeping: a verdict that arrives after the prospect has hung up is not a verdict, it is an autopsy.