Home / Tech / Google says its latest Gemini transcription model can turn your ramblings into structured text

Google says its latest Gemini transcription model can turn your ramblings into structured text

Google has revealed a new AI audio model that it says offers improved speech recognition and transcription. The company says Gemini 3.5 Transcribe is more adept at speech-to-text than earlier models, with greater precision and the ability to automatically detect more than 85 languages. It says the model can adapt a user’s unstructured speech into formatted text. You can use it to make edits with voice commands, and it removes filler words from transcribed speech too.

Google says the model can learn custom vocabulary and unique spellings, and is adept at capturing alphanumeric strings such as order numbers and postal codes. Moreover, the company claims Gemini 3.5 Transcribe can attribute speech to up to three speakers with word-level timestamps based on pre-recorded audio. That could be handy for, say, transcribing a podcast.

Check Also

Meta to pay up to $18bn to settle claims its platforms harm children

Meta Platforms will pay up to $18 billion over the next decade and strictly limit how …