← All articles
Product UpdatesOct 2, 2026 12 min read

The Accuracy Number on the Datasheet Is the Least Useful Thing on It

A bare 99% accuracy claim tells you nothing. Here is the two-question test that separates real detection specs from marketing, and what to demand instead.

The Accuracy Number on the Datasheet Is the Least Useful Thing on It

Somewhere in the evaluation folder on your desk is a datasheet that says "99% accurate" next to a detection product's name, and I want you to throw that line out. Not because the number is necessarily a lie. Because you cannot do anything with it. It is the least useful thing on the page, and the vendors who print it know that.

Here is what to demand instead: accuracy measured at decision speed, on your traffic. Anything else is a lab result. You do not run a lab. You run a call floor, and lab results do not pay agent salaries.

I have sat on both sides of this table. I bought detection tooling for outbound floors for years, and I now build it. The single biggest mistake I see buyers make is comparing two accuracy percentages as if they were the same kind of number. They almost never are. This article is the short education I wish someone had given me before my first detection purchase, written for people who sign checks, not people who read spectrograms.

Why a Bare Accuracy Claim Cannot Be Measured By Anyone

Take the claim apart. "Our detection is 99% accurate." Three things are missing, and each one of them is load-bearing.

Missing Piece One: Over How Much Audio

Was that 99% measured after listening to ten seconds of audio, or after one second?

A detector that gets to listen to ten seconds of a call before committing is playing a different sport than one that has to decide almost immediately. Any reasonable system can classify a call given ten seconds, because by then the voicemail has beeped or the human has said "hello, is anybody there" for the third time. The hard problem, the one you are actually paying to solve, is deciding correctly with almost nothing to work with.

A vendor can make any accuracy number look better simply by letting the detector wait longer before its decision is scored. Same product. Same traffic. Better datasheet. When the number does not state how much audio preceded the decision, you cannot know which sport was being played, and that is not an accident.

Missing Piece Two: At What Decision Delay

This is the piece that turns into money, so stay with me.

A detector that takes multiple seconds to be confident is not a detector you can use, no matter how confident it eventually gets. Here is what those seconds cost on a real floor. A live human answers. The detector listens. Two or three seconds tick by in silence while it thinks. The human says "hello" again, gets nothing, and hangs up. The call you paid to dial, the marketing spend that produced the number, the agent payroll waiting downstream, all of it spent on a conversation that never started, because the tool that was supposed to protect agent time instead burned the call.

There is also a compliance edge to this. Under the FTC's telemarketing rules, the abandonment rate on outbound campaigns is capped at 3 percent of answered calls per campaign per day, measured over a 30-day period. You can read the rule yourself in the Code of Federal Regulations. A slow detector that lets live answers die in silence produces exactly the kind of dead air and dropped calls that push you toward that cap. Speed is not a luxury feature. It is part of staying inside the lines.

For contrast, and because I believe in stating numbers rather than adjectives: our detection starts producing its verdict at 125 milliseconds. One-eighth of a second. That is the speed at which accuracy has to be true, or it is not true in any way that matters to your P&L.

Missing Piece Three: On Whose Calls

The third missing piece is the traffic itself. A 99% figure measured on a clean reference library of recordings is a 99% figure measured on calls that are not yours.

Real traffic is ugly. Gatekeepers answer for the decision maker. People answer in wind, in trucks, in warehouses. Voicemail greetings from some carriers are now warm, conversational, recorded by professional voice talent, and nearly indistinguishable from a hesitant human. A detector tuned on a tidy library can post a great number and then fall apart the first week it meets your actual call mix. The only accuracy that pays you back is accuracy measured on your dialing, your lists, your answering patterns.

I wrote a longer piece about how detection errors split into two directions, each with its own cost structure, here: https://amdy.io/blog/amd-error-directions. The short version: the direction of the error, whether you drop humans or send machines to agents, decides whether the error costs you revenue or payroll. A single accuracy percentage flattens that distinction, and the distinction is where all the money lives.

The Two-Question Test for Any Vendor Claim

You do not need a telecom background to pressure-test a detection claim. You need two questions, asked in order, and a willingness to accept "no" as an answer.

Question One: "At what delay was that measured?"

If the vendor cannot tell you how much audio preceded the scored decisions, the number is decoration. If they can tell you, and the answer is measured in seconds, ask what happens to the live human on the other end during those seconds. Somebody has to own that silence, and if the vendor will not, you will, in abandoned calls and dented answer rates.

The honest framing, in money terms: accuracy you can only get by waiting is accuracy you purchase with dropped calls. That is a trade, not a win. We treat the trade as unacceptable, which is why we publish the 125ms figure rather than a bigger number earned by waiting. If you want the full treatment of that trade-off, it is here: https://amdy.io/blog/amd-latency-vs-accuracy.

Question Two: "Can you show me the results on my traffic, per call, before I commit?"

This is the question that ends most vendor conversations, which is how you know it works.

A vendor confident in their detection can log every call, timestamp the verdict, and show you the per-call record: this one classified human at this millisecond mark, this one machine, this one flagged for review. That is what per-call logging means, and it is not an exotic feature. It is the difference between a claim and an observation.

If the answer is that their number comes from internal benchmarks and no, you cannot see per-call results on your own dialing during an evaluation, then you have learned everything you need. You are being asked to buy a measurement that the seller will not reproduce in your presence, on your traffic.

Detection is not a faith-based purchase category, whatever the datasheets imply. It is a measurable service running on measurable calls, and the tools to measure it exist today.

What a Meaningful Spec Sounds Like

Let me show you the shape of a spec you can actually act on, using our own numbers, because I am not going to criticize vague specs and then hand you one.

A meaningful spec states three things together, never separately. First, the decision speed: the verdict starts at 125ms from answer. Second, the verification path: every call is logged individually, so accuracy is not a vendor claim but a count you can run on your own traffic, this month and every month after. Third, the error direction is tunable and visible: you decide, knowingly, whether the system errs toward protecting conversations or protecting agent seconds, and you can see which way it erred in the logs.

That third point deserves a sentence more, because it is the one buyers miss. The expensive error is not symmetric. Dropping a live human throws away the entire cost of the call plus the revenue the conversation might have produced. Sending a voicemail to an agent wastes some seconds of payroll. One of these errors is a revenue event; the other is an efficiency event. A spec that hides the error direction is a spec that hides the majority of the financial risk. We default to the conversative side on drops, and we say so.

If you want to see how that spec translates into deployment reality, the feature set is documented at https://amdy.io/features, and installation is one command that takes about five minutes with no carrier change on your side. The evaluation does not require a migration project. It requires a trial and a look at the logs.

A Word About Compliance, Since the FTC Is Already in the Room

If you run outbound in the United States, the abandonment cap is not hypothetical. The rule, found in the Telemarketing Sales Rule as published in the Code of Federal Regulations at ecfr.gov, limits abandonment to 3 percent of answered calls, measured per campaign per day over a rolling 30-day window. Two things follow from that, and both belong in a procurement conversation about accuracy.

First, every live answer your detector drops in silence is the kind of event the cap exists to police. A detector tuned to err toward drops is not merely leaking revenue. It is manufacturing compliance surface area. When you evaluate a vendor's accuracy, you are also evaluating how much of your abandonment budget their error direction consumes, and a datasheet that says only "99%" is silent on the exact point your compliance team cares most about.

Second, verification is a compliance capability, not just a buying tool. Per-call logs with timestamps and verdicts are what you hand your counsel when a question arises about how a campaign behaved. We treat that as a first-class feature and wrote about it here: https://amdy.io/blog/amd-logs-compliance-defense. If a vendor cannot produce per-call records on demand, ask yourself what you will produce instead, and at what cost in staff hours, when the question is not academic.

How to Run the Evaluation in One Afternoon

Let me compress the whole procurement into an afternoon, because none of this requires a pilot project with a steering committee.

Before lunch, gather three numbers from your own reporting: answered calls per week, your close rate on reached humans, and the revenue per conversation those imply. You already have all three, or close enough approximations, in whatever dashboard your floor runs.

After lunch, sign up for the trial. No card is required, the Sandbox tier includes 50,000 detections per month, and installation is a single command that takes about five minutes with no change to your carrier or your dialer. Point it at a slice of real traffic, not a test list. Real traffic is the entire point.

Then read the logs. Not the marketing, the logs. Count the verdicts, look at the timestamps, and check them against calls where you know ground truth from recordings you already keep. What you are doing is replacing a datasheet claim with an observation from your own floor, which is the only comparison that was ever worth anything. If you want a worked example of the financial framing before you start, the ROI walkthrough is at https://amdy.io/roi.

That is the entire process. The vendors who make this hard are the ones whose numbers do not survive it.

The Buyer's Question Table

Here is the table I wish I had brought into my first detection vendor meeting. Left column: what to ask. Middle: what the question exposes. Right: what a defensible answer looks like. Every row reflects how we handle it ourselves, so you can hold this sheet up against any vendor, including us.

What to ask What it exposes What a defensible answer looks like
"At what decision delay was your accuracy measured?" Whether the number was earned by waiting, at the cost of dropped live answers A concrete figure in milliseconds, stated up front (ours: verdict starts at 125ms)
"How much audio preceded each scored decision?" Whether the test conditions resemble a live floor or a lab Decisions scored on the first moments of audio, the conditions your calls actually present
"Can I see per-call logs on my own traffic during a trial?" Whether the claim is reproducible or decorative Yes, with per-call logging on by default, timestamps and verdicts included
"Which direction do your errors lean, and can I choose?" Whether the vendor understands the asymmetry between lost revenue and wasted seconds A tunable bias with the trade-off stated plainly, not buried
"What does verification cost and require?" Whether "prove it" means a services project or an afternoon A free tier of 50,000 detections per month, no card required, one-command install in about five minutes
"What happens to the spec as traffic changes over months?" Whether accuracy is a snapshot or a monitored quantity Ongoing per-call logs you can audit any month, so drift shows up in your data, not in anecdotes

Print it. Bring it. A vendor who bristles at the first two rows is telling you the number on their datasheet cannot survive daylight, and you have just saved yourself a quarter of migration pain.

The trial row deserves emphasis, because it is the cheapest de-risking instrument in this entire category. Our Sandbox plan is free at 50,000 detections per month, hard-capped, no card. You can run your real traffic through it, open the logs, and count the verdicts yourself before a single dollar moves. Full pricing, including the paid tiers that start at $79 per month for 500,000 detections, is at https://amdy.io/pricing. The point of this article is not that our number beats their number. The point is that you should never again have to compare numbers that were never measured the same way.

The Deeper Lesson: Buy Measurements, Not Claims

There is a pattern across every purchasing decision in the call center stack, and detection is just the cleanest example of it. The vendors who print one impressive number and refuse to show their work are selling you the absence of information, priced attractively. The vendors who expose per-call records, decision timestamps, and error direction are selling you the ability to check up on them monthly, forever.

I know which kind of vendor I want to be stuck with after the contract is signed, because I have been the buyer on the wrong side of it. The time to find out is before signature, and the two questions above will find it out in under ten minutes.

If you take one thing from this piece, take this: an accuracy figure without a delay, without a traffic source, and without a verification path is not a specification. It is a mood. Your floor deserves specifications.

So here is my genuine question for your next vendor meeting, and it works on us too: will you show me the verdict, per call, at the moment of decision, on my traffic, before I pay? We will. The trial starts at https://amdy.io/auth/signup, and the logs will be waiting when you get there.