Your Carrier's Codec Decides Your AMD Ceiling
Every transcode between Opus, G.711 and G.729 strips acoustic detail your AMD classifier never gets back. Here is what each codec keeps and kills.
Your Carrier's Codec Decides Your AMD Ceiling
A call crosses three codecs before your AMD ever hears it. Opus out of the WebRTC SIP trunk, G.711 across the core, G.729 down the last mile to the carrier's dialer-facing leg. Each hop resamples the audio. Each hop throws away acoustic detail. By the time the stream reaches your detection logic, the model is not classifying the call that was answered. It is classifying a photocopy of a photocopy.
The industry conversation about answering machine detection treats accuracy as a model problem. Better training data, better tuning, better thresholds. I have spent years in this traffic, and I can tell you the model is only part of the equation. The other part is physics: what your carrier chain did to the audio before your classifier got a single sample. Teams will swap AMD vendors over a two-point accuracy gap while running every call through two needless transcodes, and nobody in the room asks what the audio actually looks like when it lands.
This article is about that question. What each common codec preserves, what it destroys, and why the default timing-based AMD in Asterisk degrades especially badly once compression enters the chain.
The Myth: Accuracy Lives in the Model
Here is the myth, stated plainly: AMD accuracy is a function of the classifier, so if detection is wrong, you fix the classifier.
It sounds reasonable. It is wrong in a specific, measurable way. A classifier can only operate on the features that survive into the audio it receives. If the audio arrived through a chain that low-passed it to 3.4 kHz, compressed it to 8 kbit/s, and smeared the energy envelope with codec jitter buffers, then no amount of model quality recovers what was discarded upstream. Information theory does not negotiate.
I have looked at enough detection logs to see this pattern repeat. The same campaign, the same leads, the same classifier settings, routed over two different carriers: a measurable accuracy spread between them. The model did not change. The audio did. When we built AMDY, we designed the classifier around the acoustic signature of the answer, not around a transcript, precisely because the usable signal in real carrier traffic is coarse, band-limited, and hostile. You can read more about that approach in how audio fingerprinting ML for AMD works.
The honest framing: your codec chain sets the ceiling. The model decides how close to that ceiling you get.
What Each Codec Actually Does to the Audio
Telephony codecs are standardized, public, and boring, which is exactly why people ignore them. The numbers are not opinions. They are published in ITU-T recommendations and IETF RFCs, and they describe, in effect, how much of your call's acoustic texture survives the trip.
| Codec | Bitrate | Sample rate | Band | What it preserves for AMD | What it kills for AMD |
|---|---|---|---|---|---|
| G.711 µ-law / A-law | 64 kbit/s | 8 kHz | Narrowband (300–3400 Hz) | Energy envelope, speech timing, voicemail beep spectra, silence structure | Everything above 3.4 kHz; speaker texture above narrowband |
| G.722 | 48/56/64 kbit/s | 16 kHz | Wideband (50–7000 Hz) | Fricative detail, fuller harmonics, cleaner onsets | Rare on carrier last-mile legs, so mostly academic for AMD |
| G.729 | 8 kbit/s | 8 kHz | Narrowband | Rough speech cadence, voicemail beeps (attenuated) | Fine amplitude structure; codecs add artifacts that read as noise floor energy |
| Opus (RFC 6716) | Variable, ~6–510 kbit/s | Up to 48 kHz | Fullband capable | Nearly everything, at high bitrates | Detail loss only at low bitrates; the real risk is what happens at the next transcode |
Sources: ITU-T G.711, ITU-T G.722, ITU-T G.729, and IETF RFC 6716 for Opus. These are specifications, not benchmarks.
G.711: The Honest Baseline
G.711 is companded PCM. It does not compress in the psychoacoustic sense; it quantizes logarithmically to fit 8 kHz, 16-bit-ish dynamics into 64 kbit/s. For AMD purposes it is the reference condition: what the classifier sees is close to what the channel carried. Timing is intact, the energy envelope is intact, and the spectral content above 3400 Hz was never there in the first place because the PSTN never carried it.
When teams call us and ask "what audio does AMDY actually need," the answer is G.711-class input: 8 kHz, 16-bit mono PCM. That is what our WebSocket streaming API consumes, in 20 ms chunks. Not because we dislike high fidelity, but because 8 kHz narrowband is what the world's phone network delivers, and a classifier should be built for the audio that exists.
G.729: Cheap, Everywhere, Lossy in the Wrong Way
G.729 runs at 8 kbit/s using conjugate-structure algebraic code excited linear prediction. In plain terms: it synthesizes speech from a parametric model rather than transmitting the waveform. Speech still sounds like speech, because human ears are forgiving about the exact waveform as long as the formant structure is roughly right.
AMD classifiers are less forgiving. The parametric reconstruction smooths amplitude micro-structure, and the codec's noise-fill and comfort-noise behavior can put low-level energy where silence used to be. For a classifier keying on the acoustic signature, a G.729 leg means the signature arrives blurred. It is usually still classifiable. But it is classifiable at a lower confidence margin, and lower margins are where misclassifications live.
Opus: The Best Codec That Makes Things Worse
Here is the part that surprises people. Opus is the best codec on the list, and it routinely makes AMD worse. Not through any fault of Opus, but through what happens next.
Opus is the default for WebRTC, so it dominates browser-based softphones, SIP trunking fronted by WebRTC gateways, and modern contact center stacks. The audio arrives wideband, variable bitrate, up to 48 kHz. Then it hits your dialer, which negotiates G.711. Then the carrier's last mile negotiates G.729. Two transcodes. Each one is a resample plus a re-encode, and each one applies its own band limit and its own quantization noise.
The result is audio that had pristine high-frequency content synthesized away twice, with artifacts stacked from two different codec families. The original Opus stream's quality is irrelevant. What reaches the classifier is the worst codec in the chain, plus the conversion damage from every hop before it.
Why Timing Heuristics Degrade Especially Badly
Timing-based AMD, the kind shipped in Asterisk's app_amd, does not classify sound at all. It measures durations and compares them against thresholds. If the greeting runs long, machine. If silence follows a short greeting, human. It is a stopwatch wearing a lab coat.
The defaults, for reference, from the Asterisk amd.conf documentation:
initial_silence: 2500 msgreeting: 1500 msafter_greeting_silence: 800 mstotal_analysis_time: 5000 ms
The verdict lands in the AMDSTATUS channel variable as HUMAN, MACHINE, or NOTSURE.
Now consider what a compressed, transcoded chain does to every one of those measurements.
Compressed Silence Is Not Silent
G.729 and low-bitrate Opus modes insert comfort noise during silent periods. Discontinuous transmission (DTX) means frames are not sent at all, and the decoder reconstructs an estimated noise floor. The silence detector in a timing heuristic keys on an energy threshold. Comfort noise sits near that threshold. Onset detection smears by tens of milliseconds per event.
A 2500 ms initial-silence window does not care about 30 ms of smear. But the word-gap logic does. app_amd counts words separated by pauses above 50 ms (between_words_silence) and caps at three words (maximum_number_of_words). Codec noise fill and decoder jitter buffer behavior can merge two short words into one, or split one word into two, and a heuristic counting words from an energy trace will count wrong. This is one concrete reason the default configuration drops an estimated 10–20% of live humans: the thresholds were tuned on clean G.711 test calls, and real carrier chains do not deliver clean G.711. We cover that failure mode in detail in "Hello" breaks AMD.
Jitter Buffers Rewrite Time
Every codec transition point has a jitter buffer. Jitter buffers delay, hold, and release audio in chunks to smooth network variance. From the far end the audio timing is a reconstruction, not a measurement. Two calls with identical human greetings, one over a clean channel and one over two transcodes with adaptive jitter buffers, produce measurably different pause structures at the classifier. A model that consumes the full acoustic signature can absorb some of this. A stopwatch cannot, because the stopwatch's only input is the thing that got rewritten.
What We Do Differently, and What You Should Do Regardless
I am not going to pretend the codec problem is solvable by vendors alone. It is not. But it is tractable from both ends.
On Our End
AMDY's classifier reads the acoustic signature of the answer, the energy, spectral shape, cadence, and beep structure together, rather than running a duration rulebook. It is built and trained on the audio dialers actually receive: 8 kHz narrowband, already transcoded, already degraded. The verdict starts forming at 125 ms, one eighth of a second, at 99% accuracy, and the stream protocol is 8 kHz 16-bit mono PCM in 20 ms chunks over a WebSocket. The full protocol is documented in the streaming API guide.
That design choice is a direct response to the codec reality. A model trained on clean wideband studio audio is training on audio that does not exist in production. A model trained on what carriers actually deliver has no choice but to learn the signatures that survive.
On Your End
Four things, in order of leverage:
1. Count your transcodes
Run sip show channels (or pjsip show channels on newer Asterisk), or better, capture a call leg with tcpdump and read the SDP offer/answer exchange. You are looking for the codec list on each leg. If the inbound trunk offers Opus first and your carrier leg negotiates G.729, every call crosses at least one transcode, and if your dialer's internal bridge is G.711, you have the full Opus to G.711 to G.729 stack. Most operators I ask have never looked. They know their trunk pricing to the tenth of a cent and cannot name the codec on the last mile.
The Asterisk wiki's codec transcoding documentation is the straight reference for how Asterisk selects codecs and where direct media versus transcoded paths apply. Read it once against your own dialplan and you will find at least one leg you did not know existed.
2. Prefer G.711 end to end on the AMD leg
G.711 is 64 kbit/s, so it costs bandwidth. Bandwidth is cheap. Detection errors are not. A misdetection on a live prospect costs the full acquisition spend on that lead plus the agent minute, and it does that silently, with no log line, because from the dialer's perspective nothing went wrong. If your volume makes G.729 mandatory for cost, keep the AMD analysis leg on the G.711 side of the gateway so the classifier at least sees waveform audio even if the last mile is parametric.
3. Pin codec order in your trunk config
Most transcode problems are negotiation defaults nobody chose. Explicit allow/disallow ordering in Asterisk, or the equivalent codec preference list in FreeSWITCH, removes an entire class of accidental chains. This is a five-line config change and it is the highest ratio of accuracy impact to effort anywhere in AMD tuning.
4. Audit with the audio the model saw
If you review recordings to judge AMD decisions, listen to the post-chain 8 kHz mono stream, not the pre-chain wideband capture. They are different files telling different stories; we cover exactly this trap in Detection runs on 8 kHz mono.
The Compounding Detail Nobody Tracks
One more interaction worth naming before the summary: codec damage compounds with false answer supervision. FAS is a carrier signalling a call as answered when nobody picked up, and it means your AMD is sometimes analyzing audio that is not an answer at all, ringback, comfort noise, line echo. Layer that on top of a G.729 noise floor and the classifier is asked to distinguish machine from human in a signal that is already part fiction. When you audit per-carrier accuracy, FAS and codec chain both move with the carrier, which is why carrier-level splits should be your first cut on any unexplained accuracy shift. We treat the FAS mechanics in carrier false answer supervision.
Here is the part I want operators to sit with. Every codec transition also re-clocks the audio. Sample rate conversion between 48 kHz Opus and 8 kHz G.711 involves resampling filters with their own group delay. Stack two of those plus a jitter buffer, and the timing relationship between events in the audio shifts by amounts that vary call to call.
For a human listener, irrelevant. For a transcript-based system, irrelevant. For anything measuring sub-100 ms structure in an energy envelope, it is a confound that lives in your error rate and never appears in any log. When I see a campaign whose NOTSURE share spikes after a carrier change with no config change on the dialer, this is the first thing I suspect, and it is usually the right suspect. Terminology for the rest of it is in our AMD glossary.
The Ceiling, Summed
You cannot fix your classifier past the information your carrier chain delivers. G.711 end to end gives the model everything the PSTN has to give. G.729 on the last mile shaves the margin. Opus-to-G.711-to-G.729 stacks artifacts from three encoders and hands the classifier the wreck. The model then decides how much of that ceiling you actually reach, which is why we train on degraded carrier audio rather than clean captures, and why our verdict starts at 125 ms instead of after a five-second analysis window.
Vicidial users can test this on live traffic in about five minutes with the Vicidial integration, free tier, 50,000 detections a month, no card. Compare its verdicts against your current AMD on the same calls, split by carrier, and watch the codec chain show up in your own numbers.
The takeaway line I would keep: before you blame the model, count the transcodes. If your audio crossed three codecs, part of your error rate belongs to the network, not the classifier. Track it per carrier, pin your codec order, and audit against the leg the detector actually reads. Do those three things and the remaining error rate finally means something: it is your model's true ceiling, measured on honest audio, and that number is the only one worth optimizing.
A question for peers running their own measurement: has anyone published per-codec AMD error rates from production traffic, not lab captures? I have our numbers. I have never seen a public dataset that separates G.729-leg from G.711-leg misclassification rates at dialer scale, and it is a table this industry should have.