Technology

Enterprise Voice AI Hasn't Had Its ChatGPT Moment Yet

Martin HollowayPublished 6m ago3 min readBased on 5 sources
Reading level
Enterprise Voice AI Hasn't Had Its ChatGPT Moment Yet
Image by StockSnap from Pixabay

Enterprise voice AI has not reached its "ChatGPT moment," PolyAI CTO Shawn Wen said on stage at the HumanX conference last month, despite progress on full-duplex models. TechCrunch

Wen, CTO of enterprise voice AI platform PolyAI, spoke alongside Otter CMO Alex Gay during a panel on voice interfaces. The discussion was reported on Oct. 11, 2026.

Full-duplex here means a system that can listen and talk at the same time, allowing natural interruptions. Wen said the gating problem is reasoning speed. The next challenge is making reasoning very fast so models fetch answers quickly and conversation feels natural.

For phone calls and customer-support work, that means tightly linking three steps under strict timing: finding information, reasoning about it, and turning it into speech. Latency is not cosmetic. It decides whether barge-in, backchannels and brief interruptions work as callers expect.

Wen also pointed to the front of the system. He said automatic speech recognition (ASR), the part that turns speech into text, often misses important keywords, which creates a problem in capturing full context. In production, that error spreads. A missed name or product term in the transcript becomes a missed detail for the dialogue manager and the downstream language model.

Accuracy compounds, trust erodes

Gay described the same failure from the meetings side. Gay is CMO of meeting notetaker Otter, which is working on digital twins that can stand in for people in meetings.

Otter depends directly on the base transcript. Gay said if the original transcription lacks needed accuracy, all follow-up actions become flawed and users lose trust in the platform. Summaries, action items, search and twin behavior all inherit errors in words, measured as word error rate (WER), and in speaker labels, known as diarization. Small errors compound.

That explains Otter's current focus on capture and consent mechanics. The company plans to try notifying all participants in chat that a meeting is being recorded even when its bot is not present. The proposal is framed as a fallback disclosure path for meetings where a joinable bot cannot attend or is not admitted.

Disclosure was a shared theme. Wen said it is important to establish that people are talking to AI in enterprise calls. In regulated support flows, explicit identification affects recording consent, data handling and escalation expectations, as well as basic conversational grounding.

Why the interface bet persists

The caution contrasts with longer-running predictions about voice as a primary way to use AI. ElevenLabs co-founder and CEO Mati Staniszewski says voice is becoming the next major interface for AI. TechCrunch He has separately argued that AI audio models will be commoditized over time. TechCrunch

Platform work has moved in parallel. OpenAI piloted Custom Voices, a tool that lets developers clone a voice with a 15-second sample. TechCrunch It later started rolling out advanced voice mode to some ChatGPT Plus users, a mode that allows users to speak to ChatGPT and receive real-time responses without delay and to interrupt ChatGPT during responses. Reuters

The broader context here is a split between demo readiness and deployment readiness. Full-duplex transport and interruptible output solve who talks when. They do not solve answer quality under time pressure. A voice agent must hear correctly, retrieve correctly and speak promptly, in that order, every turn.

Looking at what this means for builders, the HumanX comments point to two distinct latency budgets. One is acoustic and streaming: deciding when a speaker finished, showing partial guesses from speech recognition, and generating speech bit by bit. The other is cognitive: tool calls, knowledge lookup and reasoning steps squeezed into a few hundred milliseconds without sounding evasive. Text chat tolerates a spinning indicator. Voice does not.

In my view, the "ChatGPT moment" analogy is useful but imprecise. Chat assistants broke through when a single model capability made new behavior obvious to non-experts. Voice has no single capability to unlock. It requires joint work across speech recognition, language model reasoning and speech synthesis, plus product work on identity, consent and auditability that text largely deferred. Progress will read as reliability gains rather than a launch event.

Seen in that light, the Otter and PolyAI priorities make sense. Keyword retention, transcript fidelity and explicit AI identification are not edge concerns. They are preconditions for delegation. Users will let a twin attend or an agent resolve a billing call only if the record is correct and the counterpart is known.

Worth flagging for enterprise teams, the upside of solving those preconditions is large. Hands-free operation, shorter handling times and persistent meeting memory change how support centers and knowledge workers allocate attention. The technology arc still points toward voice as an ordinary control plane, not a novelty. It simply has not cleared the bar where ordinary users assume it will work.