Five Minutes to Install, a Week to Prove Itself
A day-by-day playbook for running a fair before/after test of AMD on your own traffic, and how to read the results honestly.
Five Minutes to Install, a Week to Prove Itself
Every vendor who has ever pitched your call center said the same thing: trust us, it works. Some showed a slide deck. A few offered a pilot that took six weeks of IT meetings and still ended with a shrug. So when I tell you that swapping your dialer's answering machine detection takes one command and about five minutes, you should not believe me either. You should test it. On your own traffic, with your own agents, against your own numbers.
This is the playbook for that test. It takes one week. It costs nothing, because the free trial runs 50,000 detections a month with no card on file. And it ends with a decision you can defend in a board meeting, not a feeling.
Why most AMD tests are worthless
I have watched a dozen call centers try to evaluate detection tools. Most of them do it wrong in one of three ways.
First, they test on a tiny sample. Two hundred calls on a Tuesday afternoon tells you nothing. Answering behavior changes by hour, by day, by list, by carrier. You need a week of normal traffic before the noise averages out.
Second, they change two things at once. New detection, new list, new script, three fresh agents. When numbers move, nobody can say why. A fair test changes exactly one variable.
Third, they measure the wrong thing. Detection accuracy sounds like the right metric until you realize you cannot independently verify it without listening to thousands of recordings. What you can measure, cleanly and cheaply, is what actually reaches your agents.
So the test below measures two numbers only: live conversations delivered per agent-hour, and agent time burned on machines. Both come out of data you already collect. Neither requires anyone's opinion.
The one-sentence tech version
AMDY's classifier runs server-side on your existing dialer, installed with a single command by your dialer admin in about five minutes, no carrier change, no hardware, and it runs alongside your current setup while you trial it. That is the whole technical burden. Everything else in this article is about running a fair test, because the install is the easy part.
Why this test is worth your Monday morning
Here is the number that should get your attention. Default dialer AMD drops an estimated 10 to 20 percent of live humans. That is not a number I found in a brochure. It is what AMDY measured across real outbound traffic: people who said hello, waited, and got hung up on by a machine that decided they were a machine.
Run the arithmetic on your own floor. Take your weekly live-answer count. Multiply by 0.15, the middle of that range. That is roughly how many humans your current setup is deleting every week. Some of them would have bought. If you want the dollar version of that math rather than doing it on a napkin, the ROI page walks through it with your own volumes.
There is a second reason that matters even if the first one does not move you. The FTC's Telemarketing Sales Rule caps predictive dialer abandonment at 3 percent of live answers per campaign over 30 days. The rule is on the books at ecfr.gov, and misclassification feeds that number. When your detection hangs up on a person, that call is an abandoned call in the regulator's arithmetic. Getting classification right is not only a revenue question. If the compliance angle is the one that keeps you up at night, I wrote more about it in the AMD logs compliance defense.
Picking the campaign to test on
Before Monday, one more decision: which campaign carries the test.
Pick the boring one. Your most stable campaign, with a list you run every week, predictable hours, and agents who have been on it for months. Stability is what makes the comparison readable. The exciting new campaign with the unproven list is the worst possible choice, because it has no baseline worth trusting and its week-to-week variance will swallow any signal.
Pick one with volume. If the campaign makes 400 dials a week, a 15 percent swing in live answers is two calls. You cannot read that. A campaign doing thousands of dials a week gives the test enough surface for the signal to clear the noise floor.
And pick one where the agents are tenured, not trainees. New agents change handle time and conversion for reasons that have nothing to do with detection, and they add variance exactly where you are trying to remove it. Tenured agents on a stable campaign, with the same supervisor all week. That is a laboratory. Everything else is weather.
Tell the supervisor, but not the agents, what is being tested. Supervisors can protect the test by resisting the urge to "fix" midweek dips. Agents, as I said, should hear nothing, because an agent who knows detection is being measured will start self-reporting weird calls, and suddenly you are running an anecdote collection instead of a trial.
The week, day by day
Monday: approve the trial
You, the owner or VP, do fifteen minutes of work today.
Sign up for the Sandbox plan. No card, 50,000 detections a month, hard cap. Nobody can accidentally spend money on a hard-capped free tier. Forward the API key to whoever administers your dialer.
Then send one email to your floor leads: next week we are running a before/after test on one metric, do not change anything else. No new scripts, no list swaps, no schedule changes during the test week. This email is the difference between data and anecdote.
Before you close the laptop, write down this week's baseline numbers, meaning the week that just ended:
- Live conversations per agent-hour (live answers your agents actually spoke with, divided by agent hours logged)
- Agent minutes spent on voicemail boxes (your dialer logs disposition and handle time; machines are in there)
- Calls dropped with no disposition, the "dead air" bucket
Three numbers. Fifteen minutes. If you cannot pull the third one, do not sweat it. The first two carry the test.
Tuesday: the five-minute install
Your dialer admin runs one command on the existing server. The install script wires AMDY into the dial path and takes about five minutes. No carrier involvement, no new hardware, no dialer-script changes. If it goes sideways, you remove it and you are back where you started, because during the trial it runs alongside your current setup rather than replacing it.
The admin's whole Tuesday commitment is shorter than their coffee break. That is deliberate. Detection is a routing decision inside the dialer, and the tool that makes it should not require a migration project.
One caution: do not let Tuesday turn into a tuning day. The defaults are the test. If you start adjusting thresholds on day one, you have introduced a second variable and the week is compromised. Tuning comes after you have decided, and out-of-range tuning input is refused rather than silently clamped anyway.
Wednesday through Sunday: run it, touch nothing
The rest of the week is discipline, not work.
Same lists, same hours, same agents, same script. Do not tell the agents what changed, because agents who know they are being measured perform differently, and that contamination ruins your comparison. If your floor runs multiple campaigns, pick your most stable one for the test and let the rest run as-is.
Watch only for operational problems, not results. Day one numbers are noise. I have seen owners panic over a bad Wednesday morning and kill a test that would have shown a clean win by Friday. Daily volatility in outbound is normal. That is exactly why the test runs a full week and not a day.
If something breaks, you will know within the hour, because agents stop getting calls or start getting garbage. The trial's per-call logging means every classification is recorded, so if a question comes up about a specific call you can look at what the classifier actually decided. You are never debugging a black box.
The comparison, done honestly
On Monday morning of the following week, pull the same three numbers for the test week and put them next to the baseline. Here is where honesty matters, so three rules.
Rule one: use full weeks on both sides. Not your best baseline week against the test week. If last week had a holiday, reach one week further back.
Rule two: normalize by agent-hours. A week where you added five agents will show more conversations and prove nothing. Conversations per agent-hour is the number that survives staffing changes.
Rule three: if the delta is inside the noise you normally see week to week, say so. If your live-answer rate normally swings 8 percent between random weeks, a 4 percent improvement is not a finding. A fair test can return a no. The point of running one is that you respect the answer either way.
The one-week test plan
| Day | Who | What | Success looks like |
|---|---|---|---|
| Monday | Owner / VP | Approve trial, sign up Sandbox (no card, 50,000 detections/mo), freeze all other changes | Baseline numbers written down |
| Tuesday | Dialer admin | Run the one-command install, ~5 minutes, no carrier or hardware change | Calls flowing normally by first break |
| Wed to Sun | Floor | Run normal traffic unchanged, agents not told what changed | No operational complaints |
| Monday (day 8) | Owner / VP | Pull same metrics for test week, compare per agent-hour | A clear yes, a clear no, or a redo with a longer window |
What not to measure
Just as important, because these are the metrics people reach for first and they will waste your week.
Do not measure connect rate. Connect rate is answered calls over dialed calls. It moves with list quality, time of day, and carrier behavior. Detection does not change how many people answer. It changes what happens after they answer. If your test shows a connect-rate change, something else moved, and you should check what before trusting anything else in the week.
Do not measure detection accuracy directly. You would need a human to listen to a sample of recordings and hand-label each one as human or machine, then compare against the classifier's verdict. It is a legitimate method and we document a version of it ourselves, but for a one-week owner test it is heavy, slow, and easy to do wrong with a small sample. The downstream numbers carry the verdict. Accuracy is the vendor's problem to prove, conversations delivered is your problem to count.
Do not measure revenue during the test week. One week is too short for conversion noise to average out, and small-sample conversion rates will lie to you with total confidence. I have seen a test week show a 30 percent conversion spike on eleven sales and an owner nearly reorganized the floor around it. Revenue enters later, in the projection, where it belongs.
Do not survey the agents midweek. "Seems better" is not a metric, and asking for it mid-test contaminates the observation. If you want agent impressions, collect them after day seven, for color, never for the decision.
Reading the results
Three outcomes are possible and only one of them is a decision.
If live conversations per agent-hour rose by anything near the 10 to 20 percent the misclassification data implies, you have your answer. Do the money math: extra conversations times your historical conversion rate times average deal size. Compare that against the pricing page, where the paid tier that fits 500,000 detections a month costs $79. I have never seen a floor where the recovered conversations did not cover that by an order of magnitude, but run your own numbers, not my confidence.
If the numbers are flat, look at the second metric before concluding. If agent time on machines dropped while conversations held steady, your agents got hours back. That is capacity you already pay for. Whether that is worth $79 a month is a question about your labor cost per agent-hour, and the per-call logs will show you exactly where the recovered minutes went.
If the test was contaminated, a list changed, a storm took out a site, whatever, throw it out and run it again. A contaminated test is not a failed test. It is an uninformative one. The trial cap is monthly, so a rerun costs another week of patience and nothing else.
The objections I hear before the test
"It will break my dialer." It runs alongside your current setup during the trial. There is no window where you have ripped out the old thing and hope the new one works.
"My IT guy will fight it." Show him the command. It is one line on a server he already administers. The five-minute estimate is not marketing; it is the median install time.
"We tried AMD tuning before and it was a mess." Threshold tuning is exactly the wrong lever, and error direction matters more than people realize. I wrote a whole piece on error directions in AMD because most floors have been quietly tuned into the bad direction for years. A classifier that starts its verdict at 125 milliseconds, one-eighth of a second, is playing a different game than threshold guessing. The latency story is its own topic, and it is covered in latency versus accuracy in AMD.
What "runs alongside" actually means
This deserves thirty seconds, because it is the reason the risk side of the test is effectively zero and I do not want that claim taken on faith.
Your dialer already makes a routing decision on every answered call: human to an agent, machine to the drop pile. That decision is a function call inside the dialer, not a piece of carrier infrastructure. During the trial, AMDY sits on your server and takes over that one decision. Nothing about your trunks, your carriers, your numbers, or your dial plan changes. The install command wires the classifier in and points the decision at it.
Reversal is the same shape. Remove the wiring and the dialer's default detection resumes doing what it did the day before you started. There is no data migration, because the trial does not move your data anywhere to speak of; it adds a classification log next to calls you already record.
Compare that to the risk profile of the changes usually pitched at call centers: new carrier contracts, new dialer versions, hardware refreshes. Those are projects. This is a setting. The asymmetry is the whole reason a one-week, one-person test is even feasible. When the downside is "my admin spends five minutes and I spend fifteen," the question stops being whether you can afford to test and becomes why you have not.
What happens at day seven
Day seven is a decision point, and I think you should hold yourself to actually deciding. Vendors love pilots that drift into limbo because limbo costs the buyer nothing and closes nothing.
If the week said yes, pick the plan that matches your volume and move on with your quarter. If it said no, you spent zero dollars and one week, and you learned something real about your floor that your current reports were not showing you. Both outcomes beat another quarter of guessing while the default keeps dropping an estimated 10 to 20 percent of the humans who answer.
The projection, in one paragraph
Once the delta is in hand, convert it to money with your own rates, not mine. Extra conversations per week, times your conversion rate over the same window you sell in, times average revenue per deal, times the weeks in a month. That is monthly recovered revenue. Against it, cost is on the pricing page: $79 a month covers 500,000 detections, $299 covers 5 million, and every plan includes unlimited servers, so a multi-site floor does not multiply licenses. The trial's hard cap means the learning phase cannot generate a surprise invoice even if your dial rate spikes. If you would rather pressure-test the assumptions first, the ROI page does the same arithmetic with your volumes plugged in.
If you are the kind of owner who distrusts free
Fair. Free trials exist to lower the cost of saying yes, and sophisticated buyers correctly read that as a sales tactic. So invert it. Use the free week to try to find a reason to say no. Push your worst list through it, the aged one with the weird answer patterns. Watch the dead-air bucket. Have your admin pull detection logs for a handful of calls and listen to the recordings himself. The trial is per-call logged; every classification is queryable after the fact. A tool that survives your attempts to break it in week one has earned more trust than any deck.
The test is the pitch. If a vendor will not let you run one this cheap, ask yourself what they are afraid you will find.
Start the free trial on a Monday. You will know by the next one.