Home » Google Introduces New Gemini Audio Models: Gemini 3.8 Live and 3.5 Transcribe
Google Gemini 3.8 Live and 3.5 Transcribe
Technology

Google Introduces New Gemini Audio Models: Gemini 3.8 Live and 3.5 Transcribe

Google has expanded its Gemini audio lineup with new tools for real-time voice applications, announcing Gemini 3.8 Live and 3.8 Live Extended Thinking. Google made the announcement on September 15, and the models are available through the Gemini API and Google AI Studio. They target developers building voice-first applications and conversational agents.

Google also highlighted Gemini 3.5 Transcribe in the announcement as an addition to its new Gemini Audio model lineup, which launched in August. The speech-to-text model supports more than 85 languages. Gemini 3.8 Live and 3.5 Transcribe are set to enhance real-time and voice-first product experiences for developers.

As stated by Google, “New Gemini Audio models are available for developers to build more intelligent conversational experiences via the Gemini API and Google AI Studio.”

Gemini 3.8 Live Brings Real-Time Voice Intelligence

Google’s Gemini 3.8 Live addresses natural, low-latency voice interactions and can perform tasks while maintaining an ongoing conversation. The model supports asynchronous function calling for background API and tool execution. Users can continue speaking while those tasks run. It can also process live visual inputs during conversations. This allows voice agents to respond using both audio and visual context.

According to Google, Gemini 3.8 Live supports over 97 languages. It also maintains accent consistency across supported languages. The model can handle alphanumeric information such as confirmation codes and claim numbers. It can also merge live audio with structured data.

Google introduced Gemini 3.8 Live Extended Thinking for more complex audio tasks. It performs deeper reasoning while maintaining the conversation. The model can provide progress updates while completing multi-step tasks. This approach can reduce interruptions during complex voice interactions. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking have competitive pricing models, charging $0.005/min for audio input and $0.018 for audio output.

Gemini 3.5 Transcribe Targets Real-Time Speech-to-Text

Gemini 3.5 Transcribe focuses on converting speech into text with low latency. The model offers automatic language detection across more than 85 languages. It also supports speaker labeling and word-level timestamps. Its streaming version, Gemini 3.5 Transcribe Live, supports bidirectional speech-to-text streaming.

Google reports a 4.0% average Word Error Rate for streaming transcription when the non-streaming version recorded a 2.6% average Word Error Rate. The model also supports automatic code-switching during speech. Developers can add up to 1,000 custom vocabulary terms.

Its Smart Transcription mode can produce cleaner transcripts while removing disfluencies and adding structured formatting to the output.

Google Expands Its Developer Audio Stack

The new models, Gemini 3.8 Live and 3.5 Transcribe, strengthen Google’s tools for building voice-driven software. Potential applications include call center agents, captioning, and audio analytics. Developers can access Gemini 3.8 Live through the Live API. Google lists several integration partners for real-world media streaming. On the other hand, Gemini 3.5 Transcribe can process uploaded audio through the Interactions API.

Google supports audio files of up to one hour with timestamps and speaker labeling. Google’s latest releases are a significant step toward real-time voice agents. The focus now extends beyond speech generation to reasoning, task execution, and transcription.

Learn about the newest AI models and other tech breakthroughs with HiTechNectar.


Also Read:

Google Pics vs Canva: What’s Different and Which One to Choose?

Subscribe Now

    A treasure trove of valuable content and exclusive insights delivered straight to your inbox.


    Receive Updates:




    We hate spams too, you can unsubscribe at any time.