Technology
Danish Kapoor
Danish Kapoor

Google deletes unnecessary sounds from conversation with Gemini 3.5 Transcribe

Google introduced the Gemini 3.5 Transcribe model, which it developed to convert speech into text. The company describes the new model as the most accurate speech-to-text system to date. Gemini 3.5 Transcribe not only transcribes raw audio verbatim, but also converts speech into organized, formatted text. According to Google’s statement, the model was also made available to developers via Gemini API and Google AI Studio.

Gemini 3.5 Transcribe can automatically understand corrections during speech. For example, when a user says, “Let’s meet on Tuesday, no, Wednesday,” the model transfers the correct day to the final text. It also removes filler sounds like “uh” and “eee”, adds punctuation marks, and produces text that can be used directly. Special vocabulary support helps to convey technical terms, personal names and expressions with different spellings correctly.

Google states that the model automatically detects more than 85 languages. The system can also recognize regional accents and different dialects, and continues the transmission without interrupting when the language changes during the conversation. It offers speaker separation, timestamps, and word-level timestamps for up to three speakers on pre-recorded audio. Support for more than three speakers is experimental for now.

Gemini 3.5 Transcribe will also come to Chrome

According to measurements shared by Google, Gemini 3.5 Transcribe reaches an average word error rate of 4.0 percent in streaming sounds and 2.6 percent in pre-recorded sounds. In FLEURS measurement, these rates are 5.50 percent and 5.04 percent, respectively. The model prepares the final text in 70 percent less time compared to Chirp 3, introduced in 2025. Google also states that the system captures expressions consisting of letters and numbers, such as postal code and order number, more accurately in noisy environments.

The new model powers the Rambler feature available for Gboard on Android and the Gemini app for macOS. Rambler converts natural speech into regular text, while providing options to correct spelling errors and change the tone of the text with voice commands. Gemini’s macOS application, on the other hand, can summarize local files using the screen context, rearrange text between applications, and create visuals at the point where the cursor is located. The model also takes advantage of screen context and chat history in Google Antigravity’s microphone feature.

Google will add Gemini 3.5 Transcribe support to Chrome soon. Thus, users will be able to type a reply, social media post or Gemini prompt by speaking into any text field on web pages. Developers for real-time transactions gemini-3.5-transcribe-livefor recorded sounds gemini-3.5-transcribe can use the model. According to official pricing, the estimated total cost of the standard model is $0.005 per minute, while the cost of the real-time version is approximately $0.009 per minute.

TechGIndia is now on WhatsAppGet the best technology deals of the day and big news you shouldn’t miss, delivered to your phone.

Join Channel

Danish Kapoor