Technology
Danish Kapoor
Danish Kapoor

Gemini 3.8 Live has arrived: Google opens a new era in voice calling

Google focuses on voice AI conversations Gemini 3.8 Live with Gemini 3.8 Live Extended Thinking introduced its models. The official announcement says that the first model is aimed at smooth and low-latency conversations, while the second model is aimed at tasks that require more steps. Both options advance the idea of ​​adapting to a new request mid-conversation, assessing visual context, and using tools in the background. However, the extent to which these capabilities are unlocked in which product depends on the version used.

Gemini 3.8 Live can process audio, video and text inputs together. Google states that the model detects switching between 97 supported languages ​​during a conversation. This can reduce the need to log in separately when meeting with a group of speakers of different languages ​​or continuing a topic in another language. The almost real-time image interpretation feature aims to make it easier to ask questions about the object or environment via the phone camera.

The Extended Thinking option comes into play in more complex workflows. According to Google, the model can continue to talk to the user while executing multi-step thinking and tool calls in the background. In this way, the system can give short feedback about progress instead of completely silent and preparing a long response. This approach is especially thought out for tasks such as booking, working on documents or collecting information from several services.

Google also shares that Extended Thinking gets high results in some voice task assessments. The Speech to Speech Quality Index score quoted by the company is 82.6. Although such measurements are useful for comparing models under the same test conditions, they do not tell the whole of daily use. Turkish speech quality, ambient noise, internet connection and limits of connected vehicles will also change the result.

Where can Gemini 3.8 Live be used?

The Google DeepMind model card shows that both versions can receive audio, images, video and text; It lists that it supports context for up to 128k tokens and output for up to 64k tokens. The card also makes it clear that limitations such as misinformation, slowdowns and timeouts may persist. Especially in an important transaction, the smoothness of the voice response does not eliminate the need for additional verification of the content.

On the distribution side, Gemini 3.8 Live; The Gemini API is included in AI Studio, the Gemini app, and some of Google’s cloud services. API and Gemini application are also specified for Extended Thinking; Using Docs Live, Gmail Live, and Keep Live within Google Workspace is also mentioned. Access scope may vary by account, region, and product tier. Therefore, we should not expect all the examples in the announcement to appear immediately for every user.

The point that Google highlights this time is not only to produce more natural sound, but also to be able to advance the work while the conversation continues. The ability for developers to build a voice interface using tools via the Live API is expanding. For the end user, the real test will be how consistently these complex flows will work in Turkish and in everyday conditions. The company published performance claims; Independent usage experience will provide a clearer picture over time.

Danish Kapoor