← All articles
Product UpdatesSep 28, 2026 13 min read

Your AMD was tuned for a phone network that no longer exists

Answer-side audio changes every year: new voicemail platforms, call screening, new codecs. Static AMD thresholds silently decay. Here is the mechanism and the fix.

Your AMD was tuned for a phone network that no longer exists

The most expensive belief in outbound calling is that answering machine detection accuracy is a fixed property of the product you bought. It is not. It is a property of the relationship between your detector and the phone network as it exists right now. The network changes. Your detector does not, unless someone retrains it.

I run AMD detection for ViciDial, Asterisk, FreeSWITCH and Issabel shops, and I can tell you the single pattern that separates operations that stay clean from operations that slowly rot: the clean ones re-measure their AMD against live traffic every quarter. The rotting ones ran one accuracy test at install time in 2023 and assumed the number held.

It did not hold. It never holds. Here is why, mechanically, and what actually decays.

The myth: accuracy is a number you bought

Every AMD vendor publishes an accuracy figure. Ours is 99% with the verdict starting at 125ms, which is 1/8 of a second from answer. That number is real, and it is measured against a specific distribution of answer-side audio: the mix of humans, answering machines, carrier announcements, SIP error prompts and dead air that actually occurs on the networks our customers dial.

The mistake buyers make is treating that figure like a spec sheet entry for a router. Throughput on a router does not depend on what the internet looks like this year. AMD accuracy does. Detection is a classification problem where the input distribution is controlled by third parties: carriers, phone OEMs, voicemail platform vendors, Google and Apple. Every time one of them ships a change, your input distribution moves. In machine learning this is called model drift, and it is not a theoretical concern. It is the default failure mode of any deployed classifier.

If nobody retrains and nobody measures, the decay is silent. Your abandonment rate does not spike. Your agents just start getting a few more voicemails parked on them, and a few more humans get dropped to theAMD black hole. On a dialer doing hundreds of thousands of calls a day, a 2% drift is thousands of wrong verdicts a day, and nothing in your dashboard says "drift" anywhere.

What actually changes on the answer side

Let me walk through the concrete changes, because drift is best understood as a list of specific events, not an abstraction.

New voicemail platforms

Voicemail is not one thing. It is whatever the subscriber's carrier or phone shipped this year. The classic beep-then-message pattern that Asterisk's app_amd was designed around in the 2000s is now one variant among several. Carriers have rolled out visual voicemail where the greeting is transcribed server-side, voicemail platforms with shortened or absent beeps, and greeting prompts that start with a carrier read-out ("the subscriber you are trying to reach...") before any subscriber audio. Each of those changes the timing signature and the spectral shape of the first two seconds after answer. That is exactly the window every AMD listens to.

Google Call Screen and iOS Call Screening

This is the biggest answer-side change since voicemail-to-email. Google Call Screen on Android, and now iOS Call Screening with Apple Intelligence, put an automated voice between your dialer and the called party. The call is answered by a bot that speaks a scripted line and transcribes the caller's response. From the dialer's perspective, the call was answered, audio is flowing, and the audio is a synthesized voice asking who is calling.

A timing-based AMD has no category for this. It was tuned for human-or-machine, and this is a third thing that behaves a little like both. We wrote a full breakdown of how iOS 26 Call Screening hits predictive dialer answer rates and what we did about it: iOS 26 Call Screening: what it does to your predictive dialer answer rates. The Google side gets its own treatment in Google Call Screen detection: how AI AMD identifies intercepted calls. The short version: a model trained before these features existed will classify intercepted calls inconsistently, and your agents pay for every misclassification with a wasted connection.

Carrier screening and early media

Carriers increasingly play their own early media: STIR/SHAKEN failure prompts, spam-likelihood warnings, network announcements. Some of these are signaled as answer when no party picked up, which is false answer supervision, and FAS corrupts detection statistics because the analyzed audio is not a real answer at all. If your AMD cannot tell a carrier announcement from a subscriber, every announcement is a coin flip.

Codec and transcode rollouts

The PSTN is a chain of transcoding. A call that starts as Opus on your SIP leg may cross EVS, AMR-WB, and G.711 before it reaches a handset, and carriers update their transcode paths without telling anyone. Each codec change alters bandwidth, frame timing, and how silence is represented. G.729 comfort noise does not look like G.711 silence. A detector keyed to specific silence characteristics drifts the day a carrier in your call path changes a codec preference. We break FAS and carrier-side weirdness down in our per-carrier FAS analysis, which is one of the features we ship precisely because aggregate accuracy hides carrier-specific rot.

The table: what shifts, what it does to the audio, what static AMD does

Answer-side change What changes in the audio What static AMD does
Visual voicemail / new voicemail platforms Greeting timing shifts, beep shortened or absent, carrier read-out precedes subscriber audio Machine-detection window misses; voicetimes classified NOTSURE or HUMAN
Google Call Screen Answered by synthesized bot voice with scripted prompt, no beep, pauses structured like a human conversation Splits between HUMAN and MACHINE with no consistency; agent connected to a bot
iOS Call Screening (iOS 26) Apple screening voice answers, caller audio transcribed; timing resembles a slow human with long initial silence Often burned by initial-silence timeout; call dropped or misrouted
Carrier screening announcements / FAS Call signaled answered, early media is a carrier prompt, no party present Announcement classified as HUMAN; agent lands on nobody; stats polluted
Codec rollout (Opus/EVS/AMR-WB legs) Bandwidth and silence representation change; comfort noise replaces true silence Silence-measurement logic reads wrong durations; threshold decisions fire early or late
Enterprise IVR / auto-attendant changes Prompt re-recorded, word gaps changed, DTMF menu inserted earlier Greeting analysis lands in NOTSURE; disposition inconsistent across campaigns

Every row is a real event that happens on real networks, on someone's dialer, this month. None of them are exotic.

Why fixed thresholds cannot adapt, at all

Asterisk's built-in app_amd is timing logic. It sets AMDSTATUS to HUMAN, MACHINE or NOTSURE, and AMDCAUSE carries the threshold and duration that produced the verdict. Its defaults, from amd.conf: total_analysis_time 5000ms, initial_silence 2500ms, greeting 1500ms, after_greeting_silence 800ms. You can tune those numbers. Generations of dialer admins have spent evenings tuning those numbers.

What you cannot do is make them adapt. They are constants. They were reasonable defaults for the answer side of a decade ago, and they are the same constants today, on a network that now contains call screening bots, transcribed voicemail and three generations of codec changes. The Asterisk documentation describes the module and its parameters honestly; nothing in it claims the defaults track the network, because they do not. See for yourself: Asterisk app_amd documentation.

There is a worse problem than the stale defaults. Default Asterisk AMD, as we have measured on real campaign audio, drops an estimated 10-20% of live humans. Those are people who said hello and got silence or a hangup. Every one of those is a burned lead, a potential TCPA complaint, and, under the FTC's TSR 3% abandonment discipline, pressure you do not want. Tuning the four thresholds harder to catch machines makes the human-drop problem worse. There is no constant setting of four numbers that solves a shifting distribution. It is the wrong tool shape.

A trained model is the right tool shape, because a model can be retrained. When iOS Call Screening shipped, we did not push new constants to a config file. The classification pipeline got new labeled examples of intercepted calls, retrained, and the verdicts changed for calls that had previously been ambiguous. That is the entire difference: drift is fixable in a model and permanent in a config file.

How drift actually shows up in your numbers

Nobody sends you a memo that your AMD drifted. You see it in second-order metrics, which is why measurement discipline matters more than any vendor claim. What I tell every operator to watch:

Human-connect rate by week

If your human-connect rate sags while contact rates hold, either your data got worse or your AMD is eating humans. Segment it: by carrier, by destination state, by campaign. Drift from a carrier codec change concentrates on one carrier. Drift from a call-screening rollout concentrates on one handset ecosystem's carriers.

NOTSURE rate

In Asterisk terms, watch the share of verdicts landing NOTSURE. In a healthy setup it is a small, stable fraction. A rising NOTSURE rate over weeks is drift announcing itself. It means the audio increasingly falls between the categories your detector knows.

Agent-verified outcomes

The only ground truth that matters is what a human heard. Log a sample of calls where the agent disposition contradicts the AMD verdict. Voicemail parked on an agent, or a human dropped after verdict MACHINE. That sample, over a few hundred calls, is your real accuracy number this month. We publish the full methodology because a vendor accuracy claim you cannot reproduce yourself is worth nothing: How to test AMD accuracy, and the ViciDial-specific audit walk-through: Measure AMD accuracy in ViciDial: the audit.

Per-carrier FAS and honeypot signals

A per-carrier FAS breakdown tells you which carriers are feeding you answered-but-nobody audio. Honeypot detection tells you when litigators and blacklist operators are planting numbers on your lists to document bad behavior. Both are early indicators that the answer side under your campaigns has changed character. Both are features we built because aggregate dashboards hide exactly this class of rot.

What retraining looks like operationally

Retraining is not a research project. It is a loop, and it runs on data you already have if you log per-call detection output. Every call that flows through our system produces a queryable detection log entry: timestamp, verdict, latency, cause, audio reference. Wrong verdicts found by agents or by recording review get labeled, labels accumulate into training data, the model gets retrained, the next week's traffic is classified by the updated model.

Recording review with classification is the quiet workhorse of that loop. Somebody, or some process, listens to a sample of calls where verdict and outcome disagree, and the label goes back into the pile. Without the recording reference, you have opinions. With it, you have training data and evidence.

This is also why latency framing matters more than people assume. A verdict that starts at 125ms is not just a UX nicety. It means the classifier commits on the first eighth of a second of answer-side audio, which means when the answer side changes, the effect on verdicts is immediate and measurable. Fast verdicts make drift visible fast. A slow AMD that waits seconds before deciding hides drift inside its own latency budget.

Drift and the compliance clock

One more angle, because it is the one that ends up costing real money. Under the FTC Telemarketing Sales Rule, 16 CFR 310.4(b)(4), your abandonment rate is capped at 3% per campaign measured over a rolling 30 days. Abandonment happens when a live human answers and no agent is there within two seconds. AMD sits directly on that path: a detector that misjudges ring timing or answer side, that burns seconds on NOTSURE, that classifies a screening bot as a machine and hangs up, manufactures abandonment. A drifted AMD is a compliance risk, not just a conversion problem. We treat the 3% rule and AMD as one subject: AMD and the TCPA 3% abandonment rule.

And TCPA record-keeping means the decay is discoverable. If a campaign is audited, the question is not what your vendor's accuracy was on a spec sheet. It is what your per-call log says happened on specific calls on specific dates, and whether your organization was measuring or flying blind.

Drift detection you can run yourself, this week

You do not need a data science team to detect drift. You need a labeled sample and the discipline to repeat the measurement. The procedure I use, trimmed to its bones:

  1. Pull a random sample of 300 to 500 connected calls from the last 14 days, stratified by carrier.
  2. For each, pull the detection verdict, the latency, and the cause, plus the agent disposition where an agent was connected.
  3. Review the recording where verdict and disposition disagree. That is your misclassification set. On a healthy system it is small enough to listen to every one.
  4. Compute the confusion counts: humans dropped, machines parked on agents, screening bots mishandled.
  5. Write the numbers down with the date. Repeat next quarter. Compare.

Step 5 is the part everyone skips, and it is the entire value. A single measurement is a snapshot; two snapshots a quarter apart are a drift estimate. If your human-drop share moved from 3% to 6% across two quarters with no change on your side, the answer side moved, and you now have the evidence to demand a fix from whoever owns your detector.

The same sample does double duty as your compliance artifact. Under the FTC's TSR, abandonment is measured per campaign over rolling 30 days, and the difference between "our vendor says 99%" and "we measured 97.2% on a stratified sample dated last week" is the difference between a marketing claim and a defensible operating record. The measurement methodology is not complicated; it is just work. Doing it quarterly is what separates an operation from an install.

Pricing for the drift era

One practical note on cost, because retraining loops are only sustainable if the detection volume is priced sanely. Our tiers: Sandbox is $0 with a hard 50K detection cap for testing the methodology above. Starter is $79/month with 500K included detections and $0.00025 per detection over. Growth is $299 with 5M and $0.00015 over. Scale is $999 with 25M and $0.00010 over. The point of that structure is that measuring, logging, and reviewing your detection should be a baseline activity, not a premium feature you buy once and abandon. The shops that treat it as baseline are the ones whose accuracy numbers still mean something in year three.

Picking a detector for a drifting world

The buying criteria fall out of everything above. Ask any AMD vendor, including us, these questions:

  1. When was the model last retrained, and against what volume of labeled calls?
  2. What happens to verdicts on a call screening bot? Not the marketing answer; the disposition code.
  3. Can I export per-call detection logs with timestamps, verdicts, latency and cause?
  4. Can I measure accuracy myself, on my traffic, this month?

If the answers are a fixed accuracy percentage and silence on the rest, you are buying a snapshot and hoping the network freezes. The network does not freeze. For a broader read on where detection stands right now, see State of AMD 2026, and for the metric set worth putting on a wall, Outbound AMD metrics to track.

The takeaway

Detection accuracy is not a number on an invoice. It is a reading of the network at a moment in time, and the moment passes. Static thresholds never adapt; a trained model can. The operators who win are not the ones who bought the best number in 2023. They are the ones who keep measuring.

Measure your AMD this quarter against live traffic, by carrier, with agent-verified outcomes. If the number you get differs from the number you were sold, you have found the drift. What you do next is the difference between a dialer operation and a settlement fund.