Buy the Trial, Not the Demo
Demos show curated audio. Trials show your traffic. A one-week evaluation design and due-diligence questions for any answering machine detection vendor.
Buy the Trial, Not the Demo
Demos show curated audio. Trials show your traffic. If you are evaluating answering machine detection this quarter, insist on the second and refuse to buy on the strength of the first.
I have sat through a lot of demos in this category. I have also run a lot of trials, ours and other people's. The gap between what a demo shows and what happens on real traffic is the single most expensive blind spot in AMD procurement. This piece is the evaluation process I would use if I were buying, written so you can run it against any vendor, including us.
Why demos are theater in this category
An AMD demo is usually a salesperson playing a handful of recordings through the classifier. The recordings are chosen. Nobody chooses a recording that makes their product look bad. The clips are short, the audio is clean, and the outcome is known before the play button is hit.
That is not lying. It is selection. But selection is exactly the failure mode you are trying to fix. Answering machine detection fails on the messy edge of your traffic: the call-center queue that opens with four seconds of hold music, the guy who answers while walking through a parking garage, the VoicemailClassFull greeting that starts with a human-sounding "hello?" before the beep. A curated demo contains none of those. Your traffic contains thousands per day.
There is a second problem. A demo tells you nothing about decision speed. #### Speed is a constraint, not a footnote
A classifier can be near-perfect if you let it listen for eight seconds. Your dialer cannot wait eight seconds. A live human who says "hello" and then hears nothing abandons the call, and abandoned calls count against you. The FTC Telemarketing Sales Rule caps predictive dialer abandonment at 3% of live answers per campaign, measured over 30 days (16 CFR 310.4(b)(4), on the books at ecfr.gov). So the only accuracy number that matters is accuracy measured at the latency your dialer actually runs at. Demos almost never state it. If you want the long version of why latency and accuracy trade off against each other, we wrote it up here: AMD latency vs. accuracy.
The short version: a demo without a stated decision time is a magic trick with the clock removed.
What a real evaluation looks like
Three ingredients. Your traffic. Your numbers. One week.
Your traffic means the actual calls your dialer places today, on your existing campaign mix, your carriers, your lead sources. Not a sample the vendor supplies. Not last year's recordings with the outcomes pre-labeled by the vendor. The whole point of the trial is to find out what the classifier does on the distribution you own, because every call center's distribution is a little different and the error rates move with it.
Your numbers means the error rates are counted by you, from your own call logs and recordings, using a method you could defend to your board. We publish a measurement procedure for exactly this reason: how to test AMD accuracy. The core of it is simple. Take a random sample of classified calls per day, listen blind, and label each one human or machine yourself. Then compare against what the classifier said. Random sampling matters. If you let the vendor pick which calls get audited, you have rebuilt the demo inside your own building.
One week because AMD performance varies by daypart and by campaign. A Monday morning sample and a Thursday evening sample see different answer behaviors. Seven days is the shortest window that averages over a full weekly cycle without turning procurement into a science project.
For scale, a useful audit sample size is:
Sample per day = 100 to 200 randomly selected answered calls
At 150 calls a day for seven days you get roughly 1,000 audited calls. That is enough to see a 3 to 4 percentage point difference in live-human error rate with a straight face. It is about ten hours of listening, which you can split across two QA people.
The two numbers that decide the purchase
From your audit you compute two rates. Everything else is decoration.
Live-human drop rate. Of the calls a human actually answered, the share the classifier called a machine and your dialer hung up on or dumped to voicemail treatment. This is the number that costs you revenue. The default AMD built into Vicidial-class dialers drops an estimated 10 to 20% of live humans; we measured that across detection data and it is the reason AMDY exists (features). Whatever vendor you evaluate, this is the before-and-after number.
Machine pass-through rate. Of the calls that were actually machines, the share the classifier called human and routed to an agent. This is the number that costs you agent minutes. It matters, but notice the asymmetry: a machine routed to an agent costs a 20-second agent hang-up. A human dropped costs the entire acquisition spend that produced the call. When a vendor quotes a single "accuracy" figure, they are blending these two errors as if they cost the same. They do not, and pretending they do is how buyers get steered toward the wrong tuning. We wrote a whole piece on which direction you want the errors to lean: AMD error directions.
The vendor questions that expose demo theater
Ask these four, in this order, and write the answers down. Vendors who survive all four are rare, which is the point.
1. "What is your accuracy measured at your production decision time, on traffic like mine?"
Not on a benchmark. Not on their lab set. At the latency they run in production, which for us is a verdict starting at 125 milliseconds, on calls from a vertical and lead mix like yours. If they answer with a number but no latency, or a latency but no number, the demo is still running. A vendor who says "it depends, run the trial and measure" is being honest. That is a better answer than a confident 98%.
2. "Can I audit the per-call log myself, call by call, during the trial?"
If the classification log is a dashboard aggregate you cannot drill into, you cannot verify anything. Per-call logging with export is the minimum evidence standard. It is also, incidentally, what your lawyers will want later; there is a compliance angle to keeping detection logs that holds up under dispute, which we cover in AMD logs as compliance defense. A vendor who cannot show you the individual call, its verdict, its latency, and the audio it based the verdict on is asking you to trust the dashboard. Don't.
3. "What happens to the borderline calls?"
Every classifier has an uncertain bucket. Asterisk's own app_amd literally returns a third verdict, NOTSURE, when nothing crossed a threshold cleanly. Ask what the vendor does with theirs: does an uncertain call go to the agent (safe, costs agent time) or to the hang-up path (cheap, costs you humans)? There is a right answer for your economics and a wrong one, but "we don't have borderline calls" is not an answer at all.
4. "How does performance drift, and how would I know?"
Models degrade as carrier behavior, device mixes, and greeting customs shift. Ask what monitoring exists and what the vendor does when drift shows up. We treat this seriously enough to have written about the failure mode openly: AMD model drift. A vendor with no drift story today has an incident in your future.
The due-diligence table
Here is the checklist I would hand a procurement lead, with what to demand and what each row is really testing.
| What to demand | What "yes" looks like | What it actually tests |
|---|---|---|
| Trial on your live traffic, no curated audio | Vendor provisions on your dialer or mirrors your stream within days | Whether the claim survives your call distribution |
| Accuracy stated with decision latency | A number plus a time, e.g. verdicts from 125ms, error rates on your vertical | Whether they know their own speed/accuracy trade-off |
| Per-call log export during trial | Every classification queryable: verdict, latency, audio reference | Whether results are verifiable or dashboard faith |
| Error split by direction | Separate live-human drop rate and machine pass-through rate | Whether they understand which error costs you more |
| Install effort stated up front | One command by your dialer admin, about 5 minutes, no carrier change | Whether the trial costs you engineering time it should not |
| Pricing with overage math | e.g. $79/month including 500K detections, then $0.00025 each | Whether the bill is predictable at your volume |
On the last two rows, since I work at a vendor, here are ours stated flatly so you can compare: install is one command your dialer admin runs, about five minutes, no carrier change; Sandbox is free at 50,000 detections a month with no card and a hard cap; Starter is $79 a month including 500,000 detections then $0.00025 per detection; Growth is $299 including 5 million then $0.00015; Scale is $999 including 25 million then $0.00010; servers are unlimited on all of them (pricing). If a vendor you are talking to cannot produce the equivalent in one email, that itself is a finding.
Trial design you can run against anyone
This is the same design I would use on a competitor of ours. Run it exactly like this.
Day 0: baseline without the vendor
Before any vendor touches anything, measure your current live-human drop rate using the audit method above. You need the "before" or the trial produces anecdotes instead of results. Most call centers have never measured this and the number is usually worse than anyone internally believes. Assume nothing; count.
Days 1 to 2: shadow mode
Run the new classifier in parallel, logging verdicts but not routing on them. Compare its verdicts to your incumbent's on the same calls. This costs nothing operationally and it surfaces integration problems while the stakes are zero. If a vendor cannot run shadow mode, that is a finding: it means their answer depends on controlling the routing, which is demo logic again.
Days 3 to 5: routed, low-risk campaign
Route on the new verdicts on one campaign, ideally one with meaningful volume but contained blast radius. Keep auditing 150 random calls a day, blind. Watch both error rates and watch agent-side signals: average handle time on AMD-routed calls, abandoned-short-call count, anything that tells you machines are leaking to agents or humans are vanishing.
Days 6 to 7: reconcile and compute
Close the books. Live-human drop rate before versus after. Machine pass-through before versus after. Then convert to money with your own inputs:
Monthly recovered conversations = monthly answered calls × (baseline drop rate − trial drop rate)
Monthly revenue effect = recovered conversations × your conversion rate × average value per conversion
We keep a calculator for this so nobody has to trust a vendor's version of their economics: ROI calculator, and the framing behind it in cost per number, the life metric. The honest version of this math uses your numbers, never a case study a vendor handed you. I do not want to quote you an outcome, because I have never seen your conversion rate and neither has any vendor who does quote you one.
A word on what trials cost
The reason buyers settle for demos is that trials sound expensive. In this category they are not, or should not be. A serious AMD vendor can be installed on an existing Vicidial server with one command by the person who administers your dialer, in about five minutes, without touching your carriers or rewriting dialplan. If the evaluation you are being offered requires a professional-services engagement, a carrier migration, or six weeks, the trial has been made artificially heavy. Weight is not quality. Sometimes it is just a moat.
Our version of the entry door: the free trial signup, 50,000 detections a month, no card, hard cap so the worst case is exactly $0. That is a deliberate design choice, because a trial the buyer is afraid to start is a demo with extra steps.
How to read the results without fooling yourself
Trials have their own failure modes. Three I have watched happen.
Auditor drift. Whoever listens to the sample starts optimistic on day one and rubber-stamps by day four. Counter it the way QA shops do: blind the labels. The auditor should not see what the classifier said before deciding human or machine. It is ten extra minutes of spreadsheet work a day and it removes most of the bias.
Campaign cherry-picking. If the routed trial runs only on your cleanest campaign, the numbers flatter everyone and decide nothing. Pick the campaign with your worst audio, your hardest vertical, the one your team complains about. A classifier that holds up there will hold up on the easy ones. The reverse is not true.
Small-sample certainty. A 2-point difference in drop rate measured over 200 audited calls is noise wearing a suit. Do not act on any comparison until you have at least a week of audit behind it. If two vendors are inside each other's margin, either run a second week or pick on latency, logging, and price, because the accuracy difference is not real.
One more discipline: write the decision rule before the trial starts. Something as plain as "we switch if live-human drop rate falls at least 3 points with machine pass-through up no more than 2, measured on a 7-day audit." Deciding the threshold in advance is the only protection against the post-hoc rationalization that every good salesperson is trained to provoke. Hand the rule to whoever signs the check and let them hold you to it.
Who should be in the room
The trial needs exactly three roles. The dialer admin, who runs the install and confirms routing. A QA person, who owns the daily audit sample. And the person who owns the P&L for the campaign, who set the decision rule. Salespeople from either side are welcome to observe the numbers at the end. They should not be present for the counting.
If a vendor insists on running the audit for you, or on presenting "their" analysis of your trial results as the decision artifact, you have learned something more useful than any accuracy number. You have learned how they will behave when something breaks in production.
The uncomfortable summary
Every AMD vendor, us included, has an incentive to show you the best ten seconds of audio they own. Your defense is not better skepticism during the demo. It is refusing to make the demo the decision event. Make the trial the decision event, run it on your traffic, count the calls yourself, and ask for accuracy at decision speed.
The purchase order should be signed by whoever audited 150 calls a day for a week, not by whoever watched the nicest recording.
If a vendor will not survive that week, better to find out in seven days than in the quarterly review. What did your last AMD decision get based on, a trial or a demo?