Gemini 3.5 Transcribe, the speech-to-text tool, now comes with two new capabilities: emotion detection and speaker identification. The update, which rolled out to users of the transcription product, pushes the software beyond simple word-for-word conversion into the realm of audio understanding.
What the new features do
The emotion detection layer reads vocal cues and labels them—anger, joy, stress, or other states—right inside the transcript. Speaker ID, meanwhile, distinguishes between multiple voices in a recording and tags each segment with a consistent label. The two features work together, so a call center transcript can show not just who said what, but how they said it.
For a transcription service, that's a step up from the usual output. Most tools give you text; this one gives you a rough map of the mood and the roles in the conversation. That's useful when the audio is a focus group, a sales call, a counseling session, or a legal deposition.
Industries that live on recorded conversations are the obvious beneficiaries. Customer support teams can flag angry callers without listening to every second. Researchers can scan hours of interviews for hesitation or excitement. Compliance officers can track emotional patterns in recorded statements—a different kind of audit trail.
The ability to tell voices apart also solves a practical headache. In meeting recordings or panel discussions, manually labeling speakers is tedious. Automating it saves time, and when that automation comes bundled with sentiment analysis, the transcript becomes a richer artifact than a plain text file.
That shift could affect a lot of workflows. If the tool delivers on its promise, some organizations may stop using separate tools for transcription and audio analysis, folding both into one step.
Challenging the competition
The update is clearly aimed at rivals. Most mainstream speech-to-text engines already handle punctuation, timestamps, and basic diarization. Emotion detection is rarer. By adding it, Gemini 3.5 Transcribe is carving out a niche that others haven't pushed hard.
It also puts pressure on smaller vendors who specialize in sentiment analysis of audio. They now face a broader platform that offers the feature as part of its core transcription service, not an add-on. How competitors respond—whether they match the feature or pivot to something else—will be a story to watch.
The intent behind the update
The stated aim is simple: to enhance the user experience. But the broader goal is to make the tool indispensable. Transcription is a commodity; emotion detection is not. By folding these capabilities into the same interface, the product becomes more than a dictation tool.
The question now is accuracy. Emotion detection is inherently messy. A single pause or a flat tone can be read differently by different listeners. The software will need to prove it can do this reliably across accents, languages, and recording quality. If it stumbles, early adopters will likely notice quickly—and competitors will have a chance to step in.



