As AI moves beyond text, voice is emerging as a key interface for real-world interactions across customer support, sales, collections and other business workflows. A 2025 Deepgram and Opus Research survey of 400 business leaders found that 67% consider voice AI core to their product or business strategy, while 84% planned to increase their voice AI budgets.
Building these applications requires more than an AI model and a phone number. Developers also need telephony connectivity, call routing, real-time audio and call controls, creating a growing infrastructure layer around programmable voice. Below are four approaches shaping the evolving infrastructure layer:
The infrastructure-first approach — As AI applications become more dependent on real-time voice interactions, some platforms are focusing on providing the underlying voice infrastructure rather than a finished application. FreJun’s Teler follows this infrastructure-first approach, providing APIs, webhooks and SDKs that allow developers and product teams to connect applications with phone networks and manage programmable voice interactions. A key component is its bidirectional WebSocket capability, which allows live call audio to be streamed to an external AI stack and responses to be streamed back into the call. Teler is also designed to be model-agnostic, allowing developers to work with their preferred large language models, speech-to-text and text-to-speech providers. The platform is initially available across India and the Middle East.
The programmable communications approach — Another approach comes from established communications platforms that provide programmable voice as part of a broader communications stack. Twilio enables developers to integrate calling capabilities into applications through APIs and SDKs. Its Programmable Voice offering covers call control, IVRs, recordings, conferencing and queues. Twilio’s Media Streams gives access to live call audio, letting developers connect voice interactions with external AI systems. This makes established programmable communications infrastructure adaptable to newer conversational use cases.
The API-led telephony approach — Plivo gives developers tools to initiate, receive and manage calls through software. Its infrastructure includes call routing, recordings and interactive voice applications. For AI use cases, live call audio streams to a developer’s own server over WebSockets, and SIP trunking connects Plivo’s phone network to AI voice agent platforms. This keeps the telephony layer separate from the AI layer, giving development teams flexibility in choosing their speech, language and AI technologies.
The real-time communications infrastructure approach — Telnyx combines voice APIs, developer tools and real-time media capabilities for applications that need to integrate communications into their products. Its infrastructure supports programmable call control, conferencing, call queuing, recording and media streaming, letting developers connect telephony with external AI systems and applications. This places real-time communications infrastructure alongside the AI layer as a core component of voice application development.
What this means for AI voice development
These approaches reflect a broader shift in how AI voice applications are being built. The AI model is only one part of the stack; production-ready applications also require reliable telephony connectivity, real-time audio, call controls and integrations with other software systems. As voice applications move from experimentation to production, developers will need to look beyond the AI model itself. Factors such as geographic coverage, latency, real-time media capabilities, developer tooling, interoperability with AI providers and the level of infrastructure control required will increasingly shape how these applications are built and scaled.
