Synthesia Moves Beyond AI Video Into Live Interactive Coaching With Roleplay Sessions

British AI startup Synthesia launched Roleplay Sessions on July 22, 2026, marking the company's first step beyond its core product of AI-generated enterprise training videos into interactive, real-time coaching. The new product puts employees in simulated high-stakes conversations — sales pitches, performance reviews, customer complaints — with an AI avatar that talks back, pushes back, and scores the participant against a rubric. TechCrunch
Roleplay is the inaugural release under a broader platform Synthesia calls Sessions. CEO and co-founder Victor Riparbelli's company plans to expand Sessions into job interviews and candidate screening, according to what TechCrunch exclusively learned. The expansion into hiring workflows would move Synthesia from training-only into the recruitment pipeline — a different buyer, a different compliance surface, and a different set of failure modes.
The two most popular use cases among early customers are sales team training and leadership or soft-skills development, particularly practicing difficult conversations. Customers can build their own training programs directly on the platform or bring in Synthesia's consultants for bespoke solutions constructed from the company's existing training documents and context.
Roleplay already has scaled commercial deployments. Early customers include one of the top three companies by market cap in Europe, one of the top five Fortune 100, and one of the biggest recruitment companies globally. The product is currently enterprise-only, but Synthesia intends to open it to small businesses, prosumers, and educational institutions within the next few months as inference costs decline.
Under the hood, the architecture is a split stack. Synthesia's proprietary technology handles avatars and voices — the visual and audio synthesis layer. The reasoning intelligence that powers the avatar's conversational behavior, its ability to push back and adapt in real time, comes from OpenAI. This is a familiar pattern in the current generation of AI products: a specialist company owns the modality-specific layer (here, synthetic video and voice) while leaning on a frontier LLM provider for the cognitive layer. It concentrates dependency risk on a single third-party API and ties Synthesia's quality ceiling to OpenAI's model improvements, but it lets a startup of Synthesia's size ship a conversational product without training its own reasoning model.
Synthesia grounds its product thesis in a meta-analysis of learning research from Rice University's Doerr Institute, which argues that practice and feedback, rather than information delivery and demonstration, are what actually drive behavior change. The framing positions Roleplay not as a novelty but as an application of established pedagogy — the variable being that the practice partner is now a synthetic agent rather than a human roleplayer or a static video branch.
Looking at what this means in practice, the shift from passive video to interactive simulation is the same transition we have watched play out across adjacent domains: chatbots moved from scripted decision trees to generative conversational agents, coding tools moved from autocomplete to agentic workflows, and training is now moving from watch-and-learn to talk-and-be-coached. The economic logic is straightforward — human roleplay coaching does not scale across an entire workforce, and a synthetic avatar that can run unlimited sessions at marginal inference cost does.
Worth flagging is the planned expansion into candidate screening. Using AI avatars to conduct job interviews raises a distinct set of concerns from training. A training simulation that scores poorly has limited downstream consequence; a screening interview that scores a candidate, and that score influences a hiring decision, enters territory already under regulatory scrutiny in the EU's AI Act and various U.S. state and local laws governing automated employment decision tools. The recruitment company already among Roleplay's early customers suggests this use case is not theoretical.
Synthesia's bet is that inference costs will fall fast enough to support broader access within months rather than years, and that enterprise demand will validate the product before it reaches smaller customers. The split-stack architecture means the reasoning quality improves with each OpenAI model generation without Synthesia doing the training itself. Whether the Sessions platform can extend from training into hiring — and whether customers and regulators will accept AI-scored interview performance — will determine whether Synthesia remains a training company or becomes something broader.


