Track Cost per Number-Life, Not Cost per Lead
Every phone number dies. The real unit cost of outbound is how many leads a number burns before spam flags kill it. Here is the math and the levers, including AMD's role.
Track Cost per Number-Life, Not Cost per Lead
A phone number you bought for two dollars can cost you forty thousand. Not in fees, in burned leads. Everyone tracks cost per lead; almost nobody tracks cost per number-life, which is the number that actually explains why identical campaigns with identical lists produce wildly different economics. I have watched shops rotate through number inventory, blame the list vendor, buy more numbers, and never once compute how many leads each number destroyed before it died.
This is the metric, and the math to compute it on your own data. I am not going to invent dollar figures. I am going to give you formulas you plug your own numbers into, because the honest version of this article is the one where you calculate it and flinch.
What a Number-Life Is
A number-life is the span between first call and spam-flagged death. Every outbound number has one. The number dials, connects, gets flagged by analytics engines, renders as "Spam Likely," answer rates collapse, and the number becomes economically dead well before anyone formally retires it.
The unit cost nobody computes is this:
cost per number-life = (number acquisition + usage-period carrying cost + all dials' termination cost) / leads contacted while the number was healthy
The denominator is the part that matters. A number is not a cost item. It is a depleting asset with a finite capacity for leads, and every bad interaction with a live human draws down that capacity faster. When people talk about spam-likely labels and caller ID reputation, they usually frame it as a branding problem. It is an asset-depletion problem.
Why Cost per Lead Hides This
Cost per lead divides spend by leads contacted. It is silent on how many leads each contact destroyed. Two shops can post the same cost per lead while one burns a number per thousand dials and the other gets five thousand healthy dials per number. Same headline metric. Five times the number inventory cost on one side, and something worse buried underneath: the second shop's later dials on a dying number are hitting prospects who see "Spam Likely," don't answer, and mentally file the caller as junk. Those prospects are not recoverable by calling again from the same number. Some are not recoverable at all.
The Funnel of a Number's Life
A number's life has stages, and each stage has a measurable rate. The table below is the audit I run when an operation tells me their numbers "die too fast." Every row is a real lever, and every lever has an owner. What it does not have is invented benchmark values, because the whole point is that your rates are yours to measure.
| Stage of number-life | What happens | What to measure | The lever that moves it |
|---|---|---|---|
| Acquisition | Number provisioned, documented on trunk | Cost per number, docs completeness | Local presence numbers, proper registration with carriers and analytics engines |
| Healthy dialing | Calls ring through, humans answer | Answer rate, connect-to-conversation rate | Attestation and reputation hygiene; see caller ID reputation scoring |
| First flags | Analytics engines begin labeling | Days to first "Spam Likely" report; complaint rate | Call frequency per number, list quality, calling hours |
| Decline | Answer rate falls as flags spread | Answer-rate slope week over week | Number rotation timing, rest cycles for flagged numbers |
| Death | Number economically unusable | Healthy dials achieved lifetime | All of the above, plus AMD accuracy in the next section |
The single most useful number in that table is healthy dials achieved lifetime. It converts number spend from a mystery line item into a depletion rate you can manage.
How AMD Misclassification Accelerates Burn
Here is where this stops being a telecom housekeeping article. Answering machine detection sits inside the number-life funnel, and misclassification is an accelerant on depletion. Both failure directions hurt, and they hurt differently.
Direction One: Dropping Live Humans
When AMD classifies a live human as MACHINE and the dialer disconnects, the prospect hears dead air or an abrupt click from an unknown number. Some fraction calls back and gets nothing. Some fraction reports the call. All of them have now had an experience indistinguishable from a robocall, because functionally it was one.
Analytics engines and carrier screening consume complaint and report signals aggressively. Ghost calls, which is what these drops are from the recipient's side, are among the fastest paths to a spam label. We treat this linkage as central: repeated ghost calls from the same number accelerate spam labeling, which then drags answer rates down on every subsequent dial from that number. One bad classification is not one wasted dial. It is one wasted dial plus damage to every future dial in that number's remaining life. That cascade is laid out in how ghost calls and spam flags destroy caller ID.
Stock Asterisk AMD is the worst offender here, and the reason is architectural, not tunable-away. The app_amd defaults: total_analysis_time 5000ms, initial_silence 2500ms, greeting 1500ms, after_greeting_silence 800ms, min_word_length 100ms, between_words_silence 50ms, maximum_number_of_words 3, silence_threshold 256. It counts silence gaps and words and returns HUMAN, MACHINE, or NOTSURE. Our measured claim is that this default setup drops an estimated 10-20% of live humans. Run the depletion math on that: at a 10% false-drop rate and a complaint-or-report conversion of even a few percent per dropped human, you are injecting flag-accelerating events into your number's life on every campaign day. The Asterisk wiki documents the module's parameters in detail if you want to check the defaults yourself: Asterisk app_amd documentation.
The Angrier Subcase
A subset of dropped humans call back. A callback into a dialer that just ghosted them hits queue logic designed for outbound, produces a second bad experience, and occasionally lands on an agent with no context. That person does not just report the number. They tell the analytics engine, their carrier, and anyone who will listen. I take callback volume after AMD-driven drops as one of the most honest internal signals of AMD health, and almost no shop measures it because the callback does not attribute cleanly to the original dial. Start measuring it. The dialer's logs have the data; the join is on calling-party number and a time window.
Direction Two: Voicemail Drops Misfired at Humans
The second failure direction is quieter and equally corrosive. When AMD classifies a live human as MACHINE and the campaign's voicemail-drop logic fires, a recorded message starts playing at a person who said hello two seconds ago. There is no faster way to get reported. From the recipient's perspective, this is the platonic ideal of a robocall: they answered, a robot talked at them. Complaint rates on misfired voicemail drops run hot enough that a small daily count can materially shorten a number's life, because the complaint signal feeds directly into the flagging engines described in the outbound spam-flag pipeline.
So both AMD failure directions deplete the same asset by different routes: drops generate ghost-call signals, misfired drops generate robocall complaints. The funnel does not care which one killed the number.
Walk the Math With Your Own Numbers
Here is the calculation I want you to run this week. No invented figures, four inputs, all from systems you already have.
Step 1: Healthy dials per number
Take numbers retired in the last quarter. For each, count dials before the first spam-flag report from your monitoring or analytics feedback source. Median across the cohort. That is your healthy-dial capacity, call it D.
Step 2: Number cost per life
Acquisition cost plus carrying cost over the observed life, divided by nothing yet. Call it N.
Step 3: Leads consumed per life
From the dialer's disposition history, count distinct leads contacted while the number was healthy. Call it L. This includes the leads you burned on bad experiences, which is the whole point, because cost per lead only counted the ones you count as worked.
Step 4: The metric
cost per number-life = (N + termination cost of all dials in the life) / L
Then run the same calculation segmenting numbers by the AMD in force during their life, if you have changed engines or tuning in the last year. The delta in L between segments, with list and campaign mix roughly constant, is your estimate of how much misclassification is costing you in number inventory. When we move a stack off stock Asterisk AMD onto our engine, which returns a verdict beginning at one-eighth of a second, 125ms, at 99% accuracy, the number-life effect is frequently larger than the immediate connect-rate effect, because the immediate effect is per-call while the number-life effect compounds over every remaining dial the number had in it.
What the Inputs Depend On
Be honest about the levers when you interpret the result.
Attestation
Numbers signed at full attestation, with identity and authorization verifiable at origination, face less screening friction, which supports answer rates and slows the flag feedback loop. Numbers on poorly documented trunks burn faster. We cover that chain in attestation and what answers your calls.
Answer rate
Every marginal point of answer rate on a healthy number draws more value from the same asset before flags land. Answer rate is downstream of reputation, which is downstream of everything in this article.
False-drop rate
Your AMD's false-drop rate sets the rate of ghost-call and complaint injection. Cut it and you slow depletion directly. This is the lever most operations never connect to number spend, because AMD config lives in the dialer and number budget lives in procurement.
Rotation Is Not a Fix, It Is a Subscription to the Problem
The standard response to number burn is rotation: buy more numbers, spread the dialing load, retire flagged ones on a schedule. Rotation works in the sense that it keeps campaigns dialing. It does nothing to the underlying depletion rate. You are renting the problem monthly instead of solving it, and the procurement line item grows quietly until someone asks why number spend tripled.
There is a worse version, too. Rotating numbers without fixing the depletion cause trains the flagging ecosystem that your number blocks, your trunk, and eventually your origination identity are a spam source. Analytics engines correlate across numbers, especially local presence blocks provisioned together. I have seen shops rotate aggressively, watch each successive cohort die faster than the last, and conclude they need to rotate even faster. That is a doom loop, and the way out of it is to attack the levers in the funnel table, not the inventory count.
When Rotation Is Legitimate
To be fair, rotation has a legitimate role. Local presence dialing, where you match area codes to the list, requires number inventory by geography, and some churn is inherent. Campaign type matters too: a compliant B2B campaign with clean lists and low frequency depletes numbers slowly enough that rotation is a rounding error. The problem is not rotation as a tactic. It is rotation as a substitute for measuring depletion. If you know your healthy-dial capacity per number, rotation becomes an engineered replenishment rate. If you do not, it is a guess that compounds.
The Rest Cycle Question
Some operators rest flagged numbers instead of retiring them. The theory is that flags decay if the number goes quiet. In practice this is uncertain territory and I would not build a plan on it: analytics engines differ, flag sources differ, and a number's recovery depends on which engines flagged it and why. What I can say with confidence is that rest cycles are worth testing against your own data, because the test is cheap. Flag a cohort, rest them 30, 60, and 90 days, and measure answer rate on reintroduction against a control cohort of fresh numbers. That experiment gives you a real number for whether rest works in your specific footprint, which beats any industry rule of thumb, including mine.
Compliance Sits Inside the Same Funnel
The FTC's Telemarketing Sales Rule puts call abandonment at a 3% cap per campaign, measured across 30 days under 16 CFR 310.4(b)(4). A dropped live human is an abandoned call. The same false-drop rate that burns your numbers inflates your abandonment ratio, so the metric is doubly load-bearing. The FTC's telemarketing rules are published directly: FTC TSR resources. We walk the abandonment mechanics in AMD and the TCPA 3% rule.
What We Built the Economics Around
Our pricing is per detection with unlimited servers on every tier, from Sandbox at zero dollars with a 50K monthly cap, through Starter, Growth, and Scale, up to Partner volumes on the pricing page. We priced it that way deliberately: if your unit cost is per detection, then the only way your detection spend rises is with volume, and the way to make the spend efficient is the same metric this article is about. A detection engine that drops 10-20% of live humans is not cheaper at any price, because you pay for it in number inventory, complaint rates, and abandoned-list capacity three line items away from the dialer config where the decision was made. The features page covers the integration paths, one line on ViciDial and Asterisk-family stacks.
Common Objections, Answered
Three objections come up every time I walk an operation through this metric.
"Our numbers are free with the trunk." No number is free. If the carrier bundles inventory, the depletion cost surfaces as termination spend on dials nobody answers and list capacity burned on prospects who filed you as spam. The asset is depleting whether or not the invoice line exists.
"We track answer rate already, isn't that enough." Answer rate is a symptom, measured per dial. Number-life is the asset, measured across dials. Two operations with identical answer rate averages can have wildly different depletion if one's rate is stable and the other's collapses in week two of every number's life. The average hides exactly the information that matters.
"Our AMD is fine, we tuned it." Tuning stock Asterisk AMD improves it within the ceiling of a silence detector, and the community-documented ceiling for that architecture is well below what a modern engine achieves. More to the point here: even a modest false-drop rate, the thing tuning attacks, is only one of the depletion levers. Attestation hygiene, list quality, and calling frequency all draw down the same asset. The metric exists to tell you which lever is doing the damage, and no single dialer setting can answer that.
A Starting Template
If you want the minimum viable version of this discipline, it is three numbers on a weekly cadence. Median healthy dials per retired number. Distinct leads contacted per retired number. And complaint or report count per number, from whatever feedback source you have, analytics engine portals or carrier reports. Plot all three over eight weeks and segment by whatever changed in your operation, trunk, tuning, list source. The segments will do the arguing for you.
The first time most shops run this, the discovery is not subtle. Either numbers are dying far faster than anyone assumed, or the leads-per-number number is embarrassingly small next to the number count in service. Both discoveries are cheap at eight weeks of attention and expensive at a year of ignorance.
The Takeaway
Cost per lead tells you what a contact cost you. Cost per number-life tells you what your operation's behavior cost the asset the contact rode on. Compute it once, segment by AMD and by attestation hygiene, and you will know exactly which lever is burning your numbers.