Technology

Modulate Raises $25M to Verify Voice Agents for Scams, Fakes and Rule-Breaking

Martin HollowayPublished 6d ago3 min readBased on 2 sources
Reading level
Modulate Raises $25M to Verify Voice Agents for Scams, Fakes and Rule-Breaking
source:modulate.ai

Modulate has raised $25 million for its voice models and analysis suite. The Boston-based voice intelligence startup closed a round led by Future Ventures, with participation from Hyperplane and Lakestar TechCrunch. PitchBook data cited in the announcement puts prior fundraising at $41 million on a $170 million valuation.

Modulate was founded in 2017 by Mike Pappas and Carter Huffman, who met as MIT physics undergraduates. The company started with voice modulation, software that changes how a voice sounds, for gaming.

The current platform centers on small models for transcription, emotional analysis, deepfake and AI music detection, and policy enforcement for voice agents in regulated industries. Transcription means turning speech into text. A small model is a compact AI trained for one narrow task, and policy enforcement means checking that automated calls follow legal and company rules. The stack runs more than 100 models split between signal extraction and analysis and detection. Signal extraction means pulling measurable features from raw audio. The company specializes in deepfake detection, including alerting organizations like call centers to a possible scam, and its technology is also used to monitor cyberattacks through voice calls.

On its own site, the company describes an Ensemble Listening Model architecture that analyzes dozens of audio signals in every utterance, offered through an API with five model families behind one key on a self-serve basis Modulate. An utterance is a single spoken phrase, and an API is a standard way for other software to use a service over the internet. That catalog includes accent detection to identify speaker accent from voice signals, language detection to identify which language is being spoken, audio event detection for non-speech sounds, and voicemail detection to recognize when a call has reached voicemail or an answering machine. Modulate states its Velma Ensemble Listening Model is 2x-4x more accurate than LLMs, the large general models behind chatbots, and ranks #1 on the Hugging Face deepfake leaderboard, a public ranking of detection systems.

The company reports it has improved 563,124,894 hours of conversations and detected more than 40 million unique business risks. Headcount stands at 40 to 45 employees, with plans to add 10 more people in the coming months to bolster model building.

For teams shipping voice agents, the company describes a split. Dialogue models handle conversation. Modulate handles verification around the conversation.

In my view, that separation is what decides whether a system can be fielded in regulated deployments, where policy enforcement, scam flagging and deepfake detection cannot be bolted on after launch. Small models for signal extraction fit a different operational profile from large models, with tighter scope and clearer evaluation criteria. The funding size and hiring plan point to focus rather than scale. Ten additional model builders on a team of 40 to 45 will not transform coverage overnight, but it can deepen accuracy on narrow acoustic tasks where errors carry compliance or fraud cost. If that work holds, voice agents gain a practical layer for trust and auditability that helps regulated buyers move from pilots to production, with fewer scams getting through, cleaner transcripts and clearer accountability for every automated call.