A Reporter Built an AI Avatar to Explain One Story

A TechCrunch reporter has created an interactive digital avatar of herself that readers can converse with. The system was described on Sept. 26, with reader access provided alongside the story. TechCrunch
The scope is narrow by design. The avatar is trained on the author's story about why venture-backed startups commit more fraud than non-VC-backed startups. It will only answer questions about that story.
The build was done by Synthesia. This was the first time Synthesia made a digital avatar for a journalist, and for anyone outside Alexandru Voica. Voica is head of corporate affairs at video-generation startup Synthesia. His interactive avatar was trained to answer common press questions about Synthesia.
Creation started with numerous photos and a two-minute recording of the author's voice. Capture took place in a mini film studio inside its office. The output was not a single file. Synthesia created a personal avatar that reads a script with and without glasses, and two interactive avatars that can talk back with and without glasses.
The interactive avatar is powered by a combination of voice-to-text, video, language, and text-to-voice models. That includes Synthesia's own video and voice models. Customers can choose alternative voice models from Cartesia, ElevenLabs, Google, or OpenAI. They can host avatars on their own cloud or pay Synthesia to host them.
Synthesia was originally based in the U.K. and opened new office space in New York. It hit a $4 billion valuation earlier this year and said last year it had crossed $100 million in ARR, or annual recurring revenue. TechCrunch It also launched a product called Roleplay Sessions that lets employees practice sales pitches with an interactive AI avatar that responds and scores responses.
The broader context here is the shift from single-take synthetic video to conversational agents with a face. A script reader needs accurate likeness and prosody, meaning natural rhythm and intonation. An interactive avatar needs that plus accurate transcription under noisy input, a language model kept strictly to an approved set of material, low-latency voice synthesis, and frame-accurate rendering that keeps lips in sync. Limiting the knowledge domain to one story simplifies that containment and makes failures easier to test.
In my view, the enterprise use is the part to watch. Scored sales practice is repetitive, measurable, and tolerant of narrow domains. That is a good fit for constrained avatars. Model choice and hosting choice also fit enterprise procurement, where teams weigh voice fidelity against cost, delay, and control over audio and transcripts.
Looking at what this enables, newsroom experiments may matter less than the pattern they validate. A reporter-scoped avatar is essentially an explainer tied to a single document. The same pattern maps to product docs, compliance guides, and onboarding material. If domain guardrails hold and operating costs stay manageable, expect more organizations to offer a face for their library rather than another chatbot window.


