Technology

Claude's Voice Mode Now Supports Anthropic's Most Capable Models, Plus Gmail, Slack, and Nine New Languages

Martin HollowayPublished 2w ago4 min readBased on 1 source
Reading level
Claude's Voice Mode Now Supports Anthropic's Most Capable Models, Plus Gmail, Slack, and Nine New Languages

Anthropic has expanded Claude's voice mode to work with its Opus and Sonnet models, moving beyond the Haiku-only setup that previously defined the feature. The update, announced on July 23, 2026, also brings voice interactions into Gmail, Slack, and Canva, and adds nine new languages. The Verge

Until now, voice mode in Claude was limited to the Haiku model, the lightest of the three tiers in Anthropic's lineup. Haiku is built for speed and lower-latency inference — meaning it generates responses faster but with less depth. Opus and Sonnet sit at the higher end of the capability spectrum, handling more complex reasoning tasks. Bringing Opus and Sonnet into the voice pipeline means users can now talk to Claude's smarter models in a hands-free, conversational way that was previously text-only for those tiers.

Two mid-conversation switching capabilities accompany the model expansion. Users can toggle between text and voice within a single conversation, and they can switch between Claude models mid-thread without restarting. The practical upshot: you could start a voice exchange on Haiku for a quick lookup, then swap to Opus mid-session when the task demands deeper reasoning, without losing the conversational thread. Anthropic has not disclosed how context preservation works across model switches, so questions about how system prompts, tool-use state, and conversation history are passed between models of different sizes remain open.

The third-party integrations extend voice mode beyond Claude's own interface into Gmail, Slack, and Canva. A user could dictate a request to Claude that operates on email content, Slack messages, or Canva design files. This positioning suggests Claude's voice layer is being treated not as a standalone chat feature but as an interaction method that spans the user's existing productivity tools — the apps where work already happens.

On the language front, Anthropic has moved nine languages out of beta and into general availability for voice mode: French, German, Spanish, Hindi, Indonesian, Italian, Japanese, Korean, and Portuguese. Previously, non-English voice support existed only in beta. The shift from beta to general availability suggests Anthropic has reached a confidence threshold on multilingual speech recognition and synthesis quality across these languages, though the company has not published benchmark data for voice accuracy per language.

The breadth of this update touches three distinct axes at once: model tier, integration surface, and language coverage. Each axis addresses a different adoption barrier. Opus and Sonnet in voice mode remove the capability ceiling that Haiku-only support imposed on voice users. The Gmail, Slack, and Canva integrations embed Claude's voice layer in workflows where users already spend their working hours rather than requiring a context switch to a separate app. The nine-language expansion broadens the addressable user base well beyond English-speaking markets.

Looking at the competitive landscape, voice has become a standard feature for frontier model providers rather than a differentiator. OpenAI, Google, and now Anthropic all offer voice interaction with their top-tier models. The differentiator is shifting from whether a platform can do voice to the quality of the voice experience, the breadth of integrations, and the flexibility of the interaction model. The mid-conversation model-switching feature stands out in that regard: it treats the model as a selectable parameter within a persistent session rather than a fixed property of the conversation. Whether competing platforms match this specific capability will be worth watching.

The combination of higher-tier models, cross-application integrations, and multilingual general availability in a single update is a substantial functional expansion for a feature that, until recently, was limited to one model and one language family. The risk for Anthropic lies in execution complexity. Voice inference latency on Opus, a substantially larger model than Haiku, will be a measurable factor in user experience. Latency in voice interactions is far less forgiving than in text, where a few seconds of generation time goes unnoticed as a reader scans earlier paragraphs. A high-capability model that takes too long to begin speaking will feel less capable, not more, to the end user. How Anthropic manages inference-time tradeoffs on Opus within a real-time voice pipeline is a technical question the update raises but does not answer.