Gemini 3.5 Live Translate is Google's latest audio model purpose-built for live speech-to-speech translation, designed for developers, enterprises, and everyday users who need fluid cross-language conversations. At its core, the model delivers near real-time translation across over 70 languages while preserving the speaker's natural intonation, pacing, and pitch. This approach transforms what was once a clunky, turn-by-turn experience into a seamless dialogue that feels as natural as speaking in a single language. Available through the Gemini Live API, Google AI Studio, Google Meet, and the Google Translate app on Android and iOS, it brings advanced AI translation to billions of users worldwide. The model represents a leap forward from previous systems by generating speech continuously rather than waiting for pauses, making multilingual interactions feel immediate and human.
Traditional speech translation systems suffer from awkward pauses and delays because they wait for a speaker to finish before processing and responding. This turn-by-turn approach disrupts the natural rhythm of conversation, making it difficult to have flowing discussions across languages. Gemini 3.5 Live Translate solves this pain point with its continuous generation capability, which balances the need for context with the desire for immediacy. The model stays just a few seconds behind the speaker throughout the session, delivering fluid audio without silent gaps. For professionals in multilingual meetings, travelers navigating foreign environments, or customer service scenarios, this elimination of latency is critical. It allows participants to react in real time, ask follow-ups, and maintain engagement without the frustration of waiting for translations.
One of the standout features is continuous speech generation, which distinguishes Gemini 3.5 Live Translate from earlier turn-by-turn models. Instead of requiring the speaker to finish a full sentence, the model processes audio as it streams and generates translated speech incrementally. This technical approach means listeners hear translations with only a slight delay, closely matching the original speaker's pace. Combined with auto-language detection across 70+ languages, users do not need to manually configure source or target languages—the model identifies them automatically. This is especially valuable in scenarios where multiple languages may be spoken in the same session, such as international conferences or mixed-language family calls. The result is a natural, conversational flow that respects the rhythm of human speech.
Noise robustness is another major capability explicitly designed for real-world environments. The model can handle loud, unpredictable settings such as busy streets, crowded lobbies, or noisy call centers without significant degradation in translation quality. This is achieved through advanced audio processing that filters background sounds while focusing on the speaker's voice. For example, a traveler using Google Translate's new listening mode on Android can hold their phone to their ear during a noisy guided tour and still receive clear, translated audio directly through the earpiece. Similarly, enterprise users in open-plan offices can rely on the model during important cross-language meetings. This resilience ensures that speech-to-speech translation remains reliable wherever conversations happen.
admin
Beyond the core model, Gemini 3.5 Live Translate integrates with a range of developer platforms and enterprise tools. Through the Gemini Live API, developers can embed real-time translation into their own applications using partner integrations like Agora, Fishjam, LiveKit, Pipecat, and Vision Agents—these handle the complex media streaming infrastructure so builders focus on user experience. For content authenticity, all generated audio is watermarked with SynthID, an imperceptible marker woven directly into the output to ensure AI-generated content remains detectable. This responsible-AI feature helps prevent misinformation and is detailed in the model card. Enterprises also gain access via Google Meet, where speech translation now supports over 70 languages and 2000+ language combinations in a single meeting, far surpassing the previous five-language limit.
The overall workflow is built around a streaming architecture that processes audio in real time. When a user speaks, the model receives the audio stream and begins translating almost immediately, generating speech output that closely follows the original. It continuously balances the trade-off between waiting for more context to improve accuracy and delivering translation instantly to stay in sync. This continuous processing eliminates the start-stop nature of previous systems. The model operates without manual language selection—it automatically detects the spoken language and translates into the target language set by the user or application. This simplicity is key to its usability across Google Translate, Google Meet, and third-party apps built on the Gemini Live API, allowing a consistent translation experience regardless of platform.
Concrete use cases already in testing include Grab's integration for multilingual communication between drivers and travelers during pickups—they handle over 10 million voice calls per month. Another scenario is the Google Meet speech translation feature, enabling participants to converse across languages like English, Mandarin, and Swedish simultaneously in the same call. CJ ENM is testing the model to provide more authentic viewing experiences for global and Korean audiences, suggesting potential for entertainment dubbing and live events. For everyday users, the Google Translate app with headphones offers seamless in-person translation, while the new Android listening mode allows private translations through the phone earpiece—helpful in museums or quiet spaces. Outcomes across all cases include reduced communication friction, faster interactions, and more natural multilingual dialogues.
The target audience spans multiple segments: developers building real-time translation apps via the Gemini Live API (available in public preview in Google AI Studio), enterprise customers using Google Meet with private preview access, and general consumers through the updated Google Translate app on Android and iOS. Partner platforms like LiveKit, Agora, and Fishjam further extend reach to custom solutions. While specific pricing is not detailed in the announcement, the availability in preview tiers suggests future commercial models. The model's primary value is enabling fluid, natural speech-to-speech translation that preserves human expressiveness across more than 70 languages. By integrating across Google's ecosystem and third-party tools, Gemini 3.5 Live Translate sets a new standard for real-time voice translation, making cross-language communication as effortless as speaking a single language.
Developers building real-time translation applications via the Gemini Live API and Google AI Studio public preview. Enterprise organizations using Google Workspace with private preview access to speech translation in Google Meet. General consumers on Android and iOS using the updated Google Translate app, including travelers and professionals needing in-person or private translation through headphones or listening mode. Partner platforms (Agora, Fishjam, LiveKit, Pipecat, Vision Agents) integrate the model for custom solutions. Specific examples include ride-hailing companies like Grab facilitating multilingual driver-traveler communication, and media companies like CJ ENM exploring dubbing and localisation. The product suits anyone who needs fluid, natural speech translation across more than 70 languages in real time.
Updated 2026-06-11