Voice AI has been in pilot mode at most companies for two or three years. What is happening now is the shift from pilot to production, and the economics driving that shift are no longer subtle.
Amazon's Nova 2 Sonic is a voice model that listens, thinks, and speaks without converting anything to text in the middle. Traditional voice assistants, including the ones many businesses deployed over the past decade, worked in three separate steps: hear the words, understand the words, say something back. Each step took time and lost information. Tone, hesitation, frustration, none of that survived the process. Nova 2 Sonic handles the entire conversation in the audio domain, which is why it can pick up that a caller sounds stressed and respond with a steadier, slower voice.
The new addition is WebRTC delivery. WebRTC is the same open internet standard that makes video calls work in a browser without downloading anything. Connecting Nova 2 Sonic to it means businesses can deploy voice AI to mobile apps, smart devices, vehicles, and web interfaces using infrastructure they largely already have. The alternative, running voice over dedicated telephony pipelines, required significantly more setup. This lowers the barrier considerably.
Nova 2 Sonic also integrates directly with Amazon Connect and major telephony providers including Twilio and Vonage. For any organisation that already runs a contact centre on one of those platforms, this is not a rip-and-replace decision. It is an add-on.
The pricing comparison is where the conversation gets serious for non-technical operators. Amazon positions Nova 2 Sonic as roughly 80% cheaper per interaction than OpenAI's equivalent real-time voice product. Independent research puts a human agent call at $7 to $12. A voice AI call on this platform runs around $0.40. That gap is not something that disappears when you factor in quality. Early adopters including customer service firm ASAPP and education company Education First have reported that the model handles noisy environments, non-native accents, and mid-sentence interruptions well enough for production use.
The language coverage is now meaningful for global operators. The model supports seven languages and features what Amazon calls polyglot voices, a single voice that can switch languages within the same conversation. A customer who starts a call in English and shifts to Hindi does not get transferred or hit a wall. That has been a real barrier for multinational deployments.
The broader market context is worth sitting with. Gartner estimates conversational AI will cut contact centre labour costs by $80 billion in 2026. That is not a prediction about what might happen. Much of it is already in motion. Twilio reported 49% year-on-year growth in voice AI revenue in 2025. The firms moving fastest are in financial services, retail, and travel, all sectors where call volume is high and query types are repetitive.
Where this leaves human agents is a real question. The honest answer is that first-line, script-driven support work is being automated at an accelerating rate. What remains for human teams is the genuinely complex, emotionally sensitive work that still goes wrong when handled by a machine. Organisations that think clearly about which calls those are, and train their teams accordingly, will be in a better position than those trying to simply hold the line.
For any operator running a customer-facing operation with more than a few hundred calls per week, the question is no longer whether voice AI will affect your cost structure. It is whether you will be the one deciding how.