Google’s Live Translate Speaks 70+ Languages As You Talk

Every second a multilingual meeting stalls waiting for a human interpreter is a second someone loses the thread entirely.

Real-time interpretation has always been the bottleneck nobody solved cheaply

Professional interpreters cost hundreds of dollars per hour and still require scheduling. Turn-based translation tools introduce lag that breaks conversation rhythm and kills the flow of a negotiation or client call.

Translated speech comes out the other side while the speaker is still talking

Google Releases Gemini 3.5 Live Translate, a Streaming Speech processes incoming audio continuously, detecting the source language automatically across 70-plus languages and outputting translated speech within seconds, preserving the original speaker’s pitch, pacing, and intonation. You pipe audio in through the Gemini Live API, Google AI Studio, or the Google Translate app on Android and iOS, and translated audio streams out in near real time. Enterprises get it through a private preview in Google Meet starting this month, while developers can hit the Gemini 3.5 Live Translate model endpoint today.

The people who feel this most work where language barriers cost money

  • Global sales teams who lose deals because a demo call in a second language kills momentum before pricing comes up
  • Developers building voice products who currently stitch together three separate APIs just to get translated speech output
  • Enterprise IT leads whose organizations run multilingual all-hands meetings where accuracy and speaker tone both matter

The distinction between who builds with this and who simply uses it will define adoption speed.

Google is betting voice AI competition moves to latency, not accuracy

Meta’s SeamlessStreaming and Microsoft’s Azure live translation both operate in this space, but Gemini 3.5 Live Translate is the first to publicly expose a single-model endpoint that handles detection, translation, and voice synthesis simultaneously without a pipeline handoff. If latency becomes the primary differentiator in live audio AI, every enterprise communication platform will face pressure to integrate or fall behind.

Four things you can build or do with this today

  • Stream a live customer support call and output translated audio for a remote agent
  • Build a multilingual voice kiosk without configuring per-language models
  • Run a multilingual team standup in Google Meet without a human interpreter
  • Test low-latency translated audio apps directly in Google AI Studio

Gemini 3.5 Live Translate is in public preview via the Live API at no listed cost for preview access, though production pricing has not been announced.

The few-second lag is real and will matter in high-stakes conversations

The model intentionally trails the speaker by a few seconds to balance context quality against speed, which means fast back-and-forth exchanges will still feel slightly off.

For teams already in the Microsoft ecosystem, Azure AI Speech with real-time translation is the closest direct alternative. For developers who want an open pipeline with more control over each stage, combining Whisper for transcription with a TTS layer remains an option, at the cost of added complexity.

The interpreter’s role in business meetings is starting to shrink

Live audio translation is crossing the threshold from novelty to infrastructure, and the tools that get there first will be embedded before anyone evaluates alternatives. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.