FreeSWITCH mod_amd vs stream audio detection: production performance
Comparing native FreeSWITCH silence detection against WebSocket audio streams across latency, false positives, and media server CPU utilization.
Split signaling, media handling and audio classification across Kamailio, FreeSWITCH and AMDY.IO, then scale each layer against its own workload.
High-throughput outbound calling is not one scaling problem. SIP signaling, media processing and answering-machine detection put different demands on a system. Put them in one tightly coupled service and a surge in one workload can complicate operations across the rest.
A practical design separates those jobs: Kamailio handles SIP routing and load distribution; a pool of FreeSWITCH servers handles call media; and AMDY.IO classifies answer audio through its WebSocket API. This is a stack pattern, not a turnkey integration. The API supports custom predictive dialer integrations, so the connector between FreeSWITCH media and the API—and the logic that acts on a result—needs to be implemented and tested for the deployment.
In this layout, the outbound dialer sends SIP calls through Kamailio. Kamailio directs new call attempts to an available FreeSWITCH node. FreeSWITCH establishes and handles the media session. The dialer or an integration service also needs to associate each call with its audio classification request, then use the resulting classification in its call flow.
That separation matters. Kamailio balances SIP signaling; it should not be treated as the component that analyzes the audio. FreeSWITCH is the media-handling layer; it does not need to own the predictive dialer’s business rules. AMDY.IO provides AI answering machine detection for FreeSWITCH and a WebSocket API for custom dialers. Its audio fingerprint analysis starts within 1/8 of a second, but that figure is not an end-to-end guarantee for a complete call flow. Network transit, media capture, integration code and dialer action all affect the time an agent waits.
For background on the media-side trade-offs, see the comparison of native FreeSWITCH detection and streamed audio detection. The important engineering question is not just how quickly a classifier returns a result. It is whether the whole path delivers the result in time for the dialer to act usefully.
Do not size this stack from a single calls-per-second target. Measure SIP attempts and failures at the routing layer, concurrent media sessions and resource use on each FreeSWITCH node, and the number and duration of active classification requests. Load-test the complete call path, including bursts and slow or unavailable dependencies. Capacity needs depend on the dial pattern, media configuration, hardware and integration behavior; there is no universal node count.
More nodes can add headroom, but they also add coordination work. Teams must maintain compatible configuration, monitor several queues and logs, and decide how calls behave when the API, a media node or the routing layer is impaired. A timeout policy that protects agent availability may classify fewer calls; waiting longer may delay an agent handoff. Measure that trade-off against the dialer’s pacing behavior. The effect of AMD latency on agent wait times is relevant when setting those limits.
Start with a limited call volume and compare detection outcomes, response time and call handling against the existing process. Log call identifiers across Kamailio, FreeSWITCH, the integration and the dialer so an operator can trace a delayed or misrouted call. Include carrier false-answer events in that review rather than counting them as ordinary AMD decisions.
The payoff of this architecture is control over independent scaling boundaries, not automatic simplicity. Kamailio can distribute signaling, FreeSWITCH can provide a media tier, and AMDY.IO can classify audio through its WebSocket API. The hard work is stitching those layers together with clear ownership, measured limits and deliberate failure behavior.
Comparing native FreeSWITCH silence detection against WebSocket audio streams across latency, false positives, and media server CPU utilization.
Long analysis windows break predictive dialer algorithms, forcing call centers to choose between agent idle time and high drop rates.
Upgrade GoAutoDial v4 audio routing to eliminate dead air and route live prospects using real-time stream analysis.