Technology

Anthropic Brings Opus and Sonnet to Claude Voice Mode, Adds Third-Party Integrations and Nine Languages

Martin HollowayPublished 2w ago4 min readBased on 1 source
Reading level
Anthropic Brings Opus and Sonnet to Claude Voice Mode, Adds Third-Party Integrations and Nine Languages

Anthropic has expanded Claude's voice mode to support its Opus and Sonnet models, moving beyond the Haiku-only configuration that previously defined the feature. The update, announced on July 23, 2026, also extends voice interactions into Gmail, Slack, and Canva, and broadens language support to nine additional languages. The Verge

Until now, voice mode in Claude was constrained to the Haiku model, the lightest of the three tiers in Anthropic's model lineup. Haiku is positioned for speed and lower-latency inference; Opus and Sonnet sit at the higher end of the capability spectrum. Bringing Opus and Sonnet into the voice pipeline means users can now interact with Claude's more capable models in a hands-free, conversational modality that was previously text-only for those tiers.

Two mid-conversation switching capabilities accompany the model expansion. Users can toggle between text and voice within a single conversation, and they can also switch between Claude models mid-thread without restarting. The practical implication is that a user could begin a voice exchange on Haiku for a quick lookup, then swap to Opus mid-session when the task demands deeper reasoning, without losing conversational context. Anthropic has not disclosed the mechanics of context preservation across model switches, so questions about how system prompts, tool-use state, and conversation history are passed between models of different sizes remain open.

The third-party integrations extend voice mode beyond Claude's own interface into Gmail, Slack, and Canva. A user could, in principle, dictate a request to Claude that operates on email content, Slack messages, or Canva design files. The integration surface suggests Claude's voice layer is being positioned not as a standalone chat feature but as an interaction modality that spans the user's existing productivity tools.

On the language front, Anthropic has moved nine languages out of beta and into general availability for voice mode: French, German, Spanish, Hindi, Indonesian, Italian, Japanese, Korean, and Portuguese. Previously, non-English voice support existed only in beta. The graduation from beta to GA suggests Anthropic has reached a confidence threshold on multilingual speech recognition and synthesis quality across these languages, though the company has not published benchmark data for voice accuracy per language.

The breadth of this update touches three distinct axes at once: model tier, integration surface, and language coverage. Each axis addresses a different adoption barrier. Opus and Sonnet in voice mode remove the capability ceiling that Haiku-only support imposed on voice users. The Gmail, Slack, and Canva integrations embed Claude's voice layer in workflows where users already spend their working hours rather than requiring a context switch to a separate app. The nine-language GA expansion broadens the addressable user base well beyond English-speaking markets.

Looking at the competitive landscape, voice has become a standard modality for frontier model providers rather than a differentiator. OpenAI, Google, and now Anthropic all offer voice interaction with their top-tier models. The differentiator is shifting from "can it do voice" to the quality of the voice experience, the breadth of integrations, and the flexibility of the interaction model. The mid-conversation model-switching feature is notable in that regard: it is an interaction primitive that treats the model as a selectable parameter within a persistent session rather than a fixed property of the conversation. Whether competing platforms match this specific capability will be worth watching.

The combination of higher-tier models, cross-application integrations, and multilingual GA in a single update is a substantial functional expansion for a feature that, until recently, was limited to one model and one language family. The risk for Anthropic is execution complexity: voice inference latency on Opus, a substantially larger model than Haiku, will be a measurable factor in user experience. Latency in voice interactions is far less forgiving than in text, where a few seconds of generation time is invisible to a reader scanning earlier paragraphs. A high-capability model that takes too long to begin speaking will feel less capable, not more, to the end user. How Anthropic manages inference-time tradeoffs on Opus within a real-time voice pipeline is a technical question the update raises but does not answer.