
Enterprise voice AI that misreads a caller’s hesitation routes them to the wrong queue, and that mistake costs a restaurant chain or insurer thousands of dropped resolutions every month.
The transcript was always the weakest link
Every cascaded voice stack strips tone, hesitation, and recognition uncertainty before the language model sees a single word. What the LLM gets is a best-guess transcript, not a conversation.
One model now handles the whole call, start to finish
PolyAI Releases Dialog takes raw caller audio as input and outputs a turn-taking decision first, before any response is generated, classifying the caller’s state as EMPTY, ONGOING, or COMPLETE in a single token. From there, the same model handles speech recognition, function calling, and response generation, delivering sub-300ms replies without running as an always-on GPU stream. TTS stays separate, so output voice stays controllable.
Call center operators are feeling this before anyone else
- Contact center directors at high-volume restaurant groups who need measurable containment gains without rebuilding their IVR stack from scratch.
- Insurance operations leads who track handle time per call and can point to the 37% latency reduction PolyAI reported in a live deployment.
- Enterprise CX architects in healthcare, hospitality, or telecom who are accountable for authentication failure rates and misdirected call volume.
Self-serve developers and SMBs are not in scope here. This is built for organizations already running hundreds of thousands of calls.
Voice AI just crossed a line the cascade model could not
Competitors like Nuance and Google CCAI still rely on multi-step pipelines where each handoff introduces latency and information loss. With PolyAI reporting 2,000 live deployments and an $86 million Series D, the pressure on cascaded architectures to justify their complexity is only going to increase.
What you can actually do with it today
- Enable audio-native dialog on existing PolyAI enterprise deployments immediately.
- Run booking, billing, and authentication flows without a separate ASR integration.
- Measure containment lift against your current cascaded baseline in production.
- Request early access if you are a new enterprise customer in restaurants, insurance, or healthcare.
Pricing is not listed publicly — contact PolyAI directly for enterprise terms.
English only at launch, with no open weights and no public API, so any team outside PolyAI’s existing customer base will need to get in line.
Alternatives
Google CCAI and Nuance mix take a cascaded approach that gives you more vendor flexibility but reintroduces the transcript bottleneck this model eliminates. If you need open-weight audio models you can self-host, Whisper-based pipelines with a separate LLM are still the only real option.
The cascade model for enterprise voice is running out of runway
Audio-native architectures are moving from research papers into live production, and the performance gap is starting to show up in real containment numbers. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.