Google launched Gemini 3.5 Live Translate today. It listens to someone speaking, translates the words into another language, and plays back the translation as natural-sounding audio, all within a few seconds. The key difference from older systems is that it does not wait for a speaker to finish a sentence before starting to translate. It runs continuously, staying a few seconds behind the speaker throughout the conversation.
The previous version of speech translation in Google Meet supported only five languages, and it could only translate to or from English. The new version covers 70+ languages and enables over 2,000 language combinations in a single meeting. That means a French, Japanese, and Portuguese speaker can each hear the conversation in their own language simultaneously, without any manual setup.
For anyone who runs international calls, customer support lines, supplier negotiations, or cross-border training sessions, this changes what is possible without a budget for professional interpreters. Grab, the Southeast Asian ride-hailing platform, is already testing the model to help drivers and passengers communicate at pickups. Their users make over 10 million voice calls per month through the app, which gives a sense of the operational scale this kind of tool can handle.
The free Google Translate app on both Android and iPhone gets the upgrade today. On Android, a new "listening mode" lets you hold your phone to your ear like a regular call and hear the translation privately through the earpiece, without headphones. That is directly useful for travel, factory floor conversations, or any situation where two people need to communicate quickly across a language gap.
For Google Workspace business customers, the rollout starts in private preview this month, with a broader release planned later this year. Pricing for that tier has not been disclosed.
The broader context matters here. The market for real-time speech translation was already crowded with specialist tools from companies like DeepL, Wordly, and Kudo, all of which built businesses around solving exactly this problem. Google embedding the same capability into Meet and into a free mobile app they already own compresses the commercial space for those standalone services. DeepL Voice, notably, still lists voice-to-voice translation as "coming soon."
There are real limits to be honest about. Automatic translation still struggles with regional dialects, industry-specific terminology, and colloquialisms. In high-stakes settings, such as legal negotiations, medical consultations, or financial discussions, errors from a generic model carry real risk. The tool is best suited to operational communication: logistics coordination, supplier calls, customer service interactions, and internal meetings where speed matters more than precision.
All audio generated by the model is marked with an invisible watermark called SynthID, which means content can be detected as AI-generated. That matters for trust and accountability as this kind of audio becomes more common in business settings.