Google's New AI Can Turn Your Voice Into Polished Text

Google launched a new speech-to-text model called Gemini 3.5 Transcribe on August 26, 2026. It takes raw audio and turns it into clean, formatted text. The model works in more than 85 languages and can recognize specialized jargon. Users can also add custom words so the model spells them correctly. Google said the model is a major step up from Chirp 3, its previous transcription model, especially when handling multiple languages and reducing errors. Google AI Blog
The model does more than just type out what you say. You can edit the text using your voice. It automatically removes filler words like "um" and "uh" and formats the output for you. If a recording has more than one person talking, the model can tell apart up to three speakers and add timestamps for each word. It is already built into several of Google's own products. 9to5Google
Google also released two companion models. Gemini 3.5 Live improves how the AI handles interruptions, detects languages automatically, and processes what it sees in real time. Gemini 3.5 Live Experimental does something different: it talks through its reasoning step by step, so users can hear how the model is thinking before it gives a final answer. The Verge
The rollout is happening in stages. The new audio features are available now in English for anyone using the Gemini app on macOS. On Android, a dictation feature called Rambler is live in select countries and languages. Chrome support is coming soon. Google said Gemini 3.5 Transcribe will eventually let users dictate text into any web field in the browser. Developers can try all three new models now through Google's AI Studio and Antigravity tools. Engadget
The broader context here is that Google has been building up a family of audio AI models for a while. In June 2026, it announced Gemini 3.5 Live Translate, which can translate spoken conversation in near real time across more than 70 languages. On the open-source side, smaller models like Gemma 3n and Gemma 4 12B are designed to run transcription directly on a laptop. Gemini 3.5 Transcribe is the most full-featured model in the group, running in the cloud.
What makes Gemini 3.5 Transcribe different from older speech recognition tools is that it handles cleanup inside the model itself. Older systems typically needed a second step to remove filler words, fix formatting, or correct mistakes. Here, those steps happen automatically. That matters for developers because it means less work to build transcription into their apps. The custom vocabulary feature helps in fields where standard speech tools often fail, like medicine, law, or product names that are hard to spell.
The ability to label different speakers and add word-level timestamps makes the model useful for more than just dictation. It can handle meetings, interviews, and podcasts where up to three people are talking. It does not yet handle large recordings with many speakers, which some businesses and media companies need.
It is worth noting that some promised features are not available yet. Dictating text into any Chrome web field would bring the model into everyday browsing, but Google has not said when that will arrive. The fact that the rollout is English-only and macOS-first suggests Google is testing the model in real-world conditions before opening it up more widely.
For developers who want to try it, the public preview in AI Studio and Antigravity is the place to start. The model is already running in Google's own products, which is a good sign, though performance on Google's own systems and performance through the developer API may not always match up, especially when many people are using it at once.


