Google has unveiled Gemini 3.5 Transcribe, a new speech-to-text model the company calls its most accurate yet. The model brings real-time streaming and multi-speaker attribution, two features that could change how live conversations get captured in text.
Real-time streaming and speaker separation
Real-time streaming means transcription can happen while audio is playing, not after the recording is done. That's a significant shift for anyone who has waited for a full recording to finish before getting a transcript. The other feature, multi-speaker attribution, lets the model tell who is talking. In a meeting transcript, that means the text is broken up by each voice, not just one continuous block of words.
Where the new features fit
Speech-to-text is already built into media tools, Google's own apps, and video editing suites. The addition of real-time output makes the model a better fit for live captioning, where the text has to keep pace with the audio. Multi-speaker attribution is useful for podcasters, journalists, and anyone handling panel discussions or interviews. For those users, the model does more than just transcribe; it structure who says what.
The accuracy claim
Google is positioning Gemini 3.5 Transcribe as its most accurate speech-to-text model. Accuracy claims are common with new models, but the combination of streaming and speaker separation suggests the company is focusing on more than just word accuracy. Handling multiple speakers in real time is a harder problem, and the model is being built to solve it.
The model's arrival doesn't change the core purpose of speech-to-text. It's still a tool for converting audio into text. The difference is that the output is now faster and better organized, which matters when transcription is used in newsrooms, courtrooms, or any other place where a conversation's structure is as important as the words themselves.
How the model handles overlapping voices in a crowded room is the test it hasn't shown yet. That's a real-world condition, and it's one the company hasn't demonstrated in its announcement.



