There are always pauses, repetitions, "uhs" and "just sayers" in speech, but Google believes that these should not be turned into text on the screen. Google has released a new generation speech recognition model, Gemini 3.5 Transcribe, which focuses on automatically converting users' "unorganized spoken language" into structured text. Filler words will be removed in the middle, and it can also understand the meaning of the speaker's self-correction in the middle, and will not copy the content before and after the change.
Officials say that this model can automatically detect more than 85 languages, with better recognition accuracy than previous versions, and can learn customized vocabulary and special spellings. It also has better extraction capabilities for mixed alphanumeric strings such as order numbers and postal codes. For pre-recorded audio files, Gemini 3.5 Transcribe can also mark the speeches of up to three speakers and attach verbatim time stamps. This feature is of practical use for organizing verbatim transcripts of podcasts.






