← All articles
Product UpdatesSep 28, 2026 13 min read

Detection Runs on 8 kHz Mono. Your Recordings Lie to You

Teams audit AMD against 16 kHz stereo WAVs and blame the model. The classifier heard 8 kHz mono PCM. Here is what Nyquist says survived.

Detection Runs on 8 kHz Mono. Your Recordings Lie to You

8,000 samples per second, 16 bits, one channel. That is the entire audio your AMD classifier evaluated. Not the 16 kHz stereo WAV you are listening to in your audit queue. Not the wideband recording your carrier dashboard streams back. 8 kHz mono PCM, downsampled through a chain of carrier equipment before it ever reached the detection engine.

Every week I see the same scene play out. A QA team reviews flagged misdetections. The analyst opens the recording, hears a bright, obvious, unmistakable live human saying "hello, thanks for calling," and files the ticket: the AMD called this a machine, the model is broken. The escalation reaches my desk with a recording attached. The recording is 16 kHz, stereo, and was captured on the agent leg after the audio had been mixed, resampled upstream of anything the classifier touched, or both. The analyst is not wrong about what they heard. They are wrong about what the model heard. Those are two different signals, and on real carrier traffic they can honestly disagree.

This article is about why. Not as an excuse for misclassifications, some are real errors, but as the correct mental model for auditing them. If your review process compares a human ear at 16 kHz against a classifier at 8 kHz mono, your review process has a sampling bias of its own.

The Myth: The Recording Is Ground Truth

The assumption underneath every one of these escalations: the recording represents what the AMD decided on.

It usually does not. In a typical ViciDial or Asterisk deployment, the detection path and the recording path diverge at the channel. The AMD engine consumes the inbound leg as delivered, 8 kHz narrowband, frequently after one or more codec transitions. The recording, meanwhile, may be mixed from both legs, converted by soxmix or a recording macro, saved post-decision at whatever sample rate the recording config specifies, and stored with zero provenance about what the classifier actually ingested.

The result is an audit performed against evidence the model never saw. I wrote previously about why AMD should classify sound, not words, and the same principle governs audits: the only valid evidence of what a classifier experienced is the exact byte stream it consumed. Anything else is a reconstruction.

Nyquist, in One Paragraph, for People Who Skipped It

The Nyquist-Shannon sampling theorem says a sampled signal can only represent frequencies below half the sampling rate. At 8 kHz sampling, the hard ceiling is 4 kHz. Not approximately, not as a guideline: everything above 4 kHz is mathematically unreachable in the signal, and in real telephony it is worse, because PSTN narrowband passes roughly 300 to 3400 Hz. The 8 kHz stream your classifier reads contains zero information above 4 kHz. Not compressed information. Zero.

Your 16 kHz recording has a ceiling of 8 kHz. It can carry more than double the frequency range of what the classifier saw. Sibilants, breath, the crisp attack of a consonant, the air in a speaker's voice, all of that lives largely above 4 kHz, all of it exists in your audit file, and none of it ever reached the model.

What 8 kHz Keeps

Here is the point that gets missed: for answering machine detection, most of what matters survives. The 300 to 3400 Hz band still carries a great deal.

  • Speech timing and energy envelope. When sound starts, when it stops, how pauses are structured across the first seconds of an answer. This is the raw material of human-versus-machine discrimination.
  • Voicemail tones and beeps. The 1000 Hz SIT-ish tones, the 800-1400 Hz beep families, and the DTMF-adjacent markers voicemail systems emit sit comfortably inside narrowband. They survive 8 kHz intact.
  • Fundamental frequency and first formants. Adult speech fundamentals sit roughly 85 to 255 Hz, and the first two formants that shape vowel identity mostly land below 3 kHz. The band-limited signal is still unmistakably speech, with intact prosody.
  • Cadence structure. The clipped, looped, or batch-generated rhythm of a voicemail greeting versus the messy onset of a live hello. This is timing, and timing transmits cleanly.

That is why AMDY's classifier consumes 8 kHz 16-bit mono PCM over a WebSocket in 20 ms chunks, with the verdict starting to form at 125 ms at 99% accuracy. The acoustic signature of an answer, energy shape, spectral balance, cadence, beep structure, lives inside what the band carries. The streaming API documentation specifies that format for exactly this reason: it is the honest format of telephony audio.

What 8 kHz Loses

  • Fricative detail. The /s/, /f/, /sh/ sounds that define consonant crispness have energy concentrated above 4 kHz. In an 8 kHz stream they are blunt shadows.
  • Speaker texture. What makes a voice sound young, hoarse, or nasal in a way your ear picks instantly is substantially high-frequency content. A classifier at 8 kHz hears the words and prosody, not the person.
  • Recording-quality cues. Room tone, mic character, background environment above 4 kHz. Gone.

For a human listener at 16 kHz, those losses are exactly what makes the "obvious human" call obvious. The analyst's certainty is partly built on evidence that was never in the classifier's possession.

Two Honest Observers, One Disagreement

Set the two views side by side. A human ear with 16 kHz stereo, full attention, and cultural priors about what voices sound like. A classifier with 8 kHz mono, one eighth of a second to begin forming a verdict, and no access to anything above 4 kHz.

When they disagree, the disagreement is not necessarily an error. It can be two accurate readings of different signals. The human says "clearly a person, I heard the voice texture." The classifier's view of the same call contains a band-limited energy trace where the onset structure looked, acoustically, like a machine pattern. Both right about their input. Only one of them made the routing decision.

That is the uncomfortable part, and I will not soften it. The classifier makes the decision on the degraded signal because that is the only signal that exists at decision time. You cannot route a call on the analyst's retrospective listen. The choice is to classify the band-limited stream well or to classify it badly, and the entire discipline of building AMD is building for the stream you actually get. We cover the mechanics of that in audio fingerprinting ML for AMD.

A Worked Example of an Honest Disagreement

Let me make the disagreement concrete, because the abstraction is where the argument gets lost. A flagged call comes in: model said MACHINE, agent never heard the call, QA says HUMAN, ticket opened.

The 16 kHz file: a woman says "hi, you've reached the front desk" with an upward inflection, there is a chair squeak at 0.8 seconds, and after she speaks you can hear an office in the background. To any listener, a person.

The decision stream, which I pulled from the log: 8 kHz mono, inbound leg only. The "hi" starts abruptly, almost clipped at the onset, which happens when a voicemail greeting was recorded too close to the mic or the answering system truncates the first syllable. The phrase runs flat in the second half, prosody that the wideband file rendered as human inflection reads as a flattened contour at narrowband. The chair squeak, most of its energy above 4 kHz, is absent entirely. What remains is a short, clipped, flat phrase with a hard onset. On those features, machine was a defensible verdict. On the wideband file, human was obviously right. Both observers were competent. They observed different signals.

I am not claiming every escalation resolves this way. Some flagged calls are flat-out model errors, and saying otherwise would be self-serving nonsense. The claim is narrower and more useful: until you have looked at the 8 kHz stream, you do not know which kind of ticket you are holding, and the two kinds go to completely different people. One goes to the model owner with PCM attached. The other goes to whoever owns your QA evidence capture, and its fix costs nothing.

There is also a training dividend. Every honest-disagreement case you reclassify from "model bug" to "information asymmetry" cleans your error dataset. Teams that label all disagreements as model failures end up fine-tuning against label noise, which is a real cost, paid in the currency of future accuracy.

Why the Recordings Specifically Lie

The word "lie" is doing precise work here. The recordings are not fake. They are misleading in three repeatable ways.

1. The Wrong Sample Rate

The most common trap, covered above. If your recording config saves 16 kHz files while your AMD consumes an 8 kHz stream, every audit is an apples-to-oranges comparison. The fix is procedural: audit against the analysis-leg audio at the analysis sample rate, or at minimum, downsample your review file to 8 kHz mono before judging the model.

2. The Wrong Leg and the Wrong Mix

Stereo recordings with caller on one channel and agent or system on the other tell a different story from the mono inbound stream. A classifier deciding at 125 ms into the answer has heard only the opening of the caller's audio. Your recording includes everything after: the agent's response, the conversation, the context. You are auditing a decision made on 125 ms against a file containing 45 seconds. The model did not have the ending when it had to choose. Hindsight is a luxury the routing decision never had.

3. The Post-Hoc Cleaner Signal

Carrier-side recordings are often captured before the last compression hop, or after enhancement, or from a monitoring tap that never fed the classifier. The transcoding story, Opus to G.711 to G.729 chains and what each stage destroys, is one we treat in depth in how codecs set your AMD ceiling. The short version for audits: if your recording came from a different point in the chain than the classifier's tap, it is evidence of a different call.

Building an Audit Process That Does Not Fool You

I run these audits myself, and the process below is what makes the results mean anything.

Capture the Decision Stream

Log the exact audio the classifier consumed. In AMDY's case the WebSocket stream is 8 kHz 16-bit mono PCM, so the audit artifact is the PCM as streamed, chunk for chunk, as specified in the streaming protocol. If you build your own tooling around an Asterisk or ViciDial integration, tap the same EAGI leg the detector reads, not the recording bridge.

What to log alongside the audio

The PCM alone is not enough to make an audit finding actionable. Attach the metadata the decision carried: the timestamp of the verdict, the chunk index it was reached on, the channel variables set at decision time, and the caller code. When an error turns out to be real, that bundle is exactly what a vendor or your own retraining pipeline needs, and when it turns out to be an information asymmetry, the timestamps are what prove it. Storage cost for a 125 ms-plus-window capture is trivial. A misrouted live prospect is not.

Why the tap point matters more than the file format

There is a temptation to keep the existing recording pipeline and just convert files. Resist it if the pipeline taps the wrong leg. The reference for how Asterisk bridges and records channels, including which legs get mixed into a recording, is the Asterisk wiki's documentation on call recording. The classifier reads the inbound leg as delivered; MixMonitor by default reads the mixed channel. Same call, different physics. Your audit tap has to match the detector's tap or you have rebuilt the original problem with extra steps.

Review at the Decision's Resolution

Downsample. Force the file to 8 kHz mono, cut it at the decision timestamp plus a small window, and listen to that. It is humbling how many "obvious human" files turn ambiguous at the fidelity the model actually had. You will find genuine errors this way too, and those findings will be actionable, because they describe the failure as the classifier experienced it.

Separate the Error Classes

Three buckets, and they need different remedies:

Bucket What it looks like in review Where the fix lives
True misclassification The 8 kHz decision stream itself shows a clear pattern the model read wrong Model side: retrain, retune, escalate to your vendor with the PCM attached
Information-asymmetry disagreement Human at 16 kHz is certain, model at 8 kHz had a genuinely ambiguous signal Process: accept, or improve tap point, but the model had no path to the answer
Wrong-audit disagreement The recording was 16 kHz, mixed, or from another chain point Audit process: fix the evidence capture, the model did nothing wrong

The middle bucket is the one that deserves more industry honesty. A model cannot classify information that never reached it, and a human auditor with double the frequency range will keep being certain in cases where certainty was not available downstream. Treating all disagreements as model bugs produces a bug queue full of unfixable entries and erodes trust in the ones that matter.

Watch the Timing on "Hello"

One more trap worth naming. A live human "hello" often lands with silence before it and a question-shaped rise after it. A voicemail greeting often starts with clipped audio or runs flat and long. At 16 kHz with human attention, the distinction feels instant. At 8 kHz with 125 ms of context, the distinction lives in onset energy and pause structure, which is exactly what our approach reads, and which is also why the word "hello" itself is a notoriously bad single-feature discriminator, as we unpack in "Hello" breaks AMD. If your audit labels every short greeting as obviously human, your audit is encoding the same heuristic that makes default timing AMD drop live people.

The Takeaway

The number to hold onto is 8,000. Eight thousand samples a second, one channel, 4 kHz ceiling. Every decision your AMD makes was made inside that budget. Audit against that budget and your error findings become signal. Audit against a 16 kHz stereo file and you are grading the model on a test it never took.

If you want to check your own stack, the fastest path is our free Sandbox, 50,000 detections a month, no card, and a streaming API that hands you the exact PCM the classifier sees. Capture it, downsample your reviews to match, and count how many of your open AMD tickets survive contact with the real signal. In my experience it is fewer than half, and the ones that survive are the ones actually worth fixing.

A genuine question for peers running QA on dialer traffic: does anyone archive the decision-stream PCM as a matter of policy rather than the mixed recording? I have never seen it as a default in any ViciDial install I have touched, and given how cheap storage is relative to a misrouted live prospect, I do not think it should be optional.