Technology

Suno Launches Speech Beta for Voiceovers With Built-In Music

Martin HollowayPublished 2d ago3 min readBased on 11 sources
Reading level
Suno Launches Speech Beta for Voiceovers With Built-In Music
source:suno.com

Suno has launched Speech, a beta feature that generates spoken voice audio from a script or a written description. The capability is available in public beta on Suno's web and mobile platforms. The Verge

Access is through the Create tab, then the Speech option. There are two input paths. Simple mode takes a short description of what you want to hear. Advanced mode takes a full custom script that is read out word for word.

Advanced settings allow control of AI voice gender, speech style, and variety of voice generation. Output can run up to about eight minutes. That length fits podcast segments, explainers, and short narration rather than brief clips.

Speech can generate a voiceover and background music together as one finished track. A toggle switches off the music for voice-only output. Chief product officer Jack Brody called it "the first audio model that generates voice and music together as one cohesive track." Suno said it will continue to improve Speech based on user feedback and acknowledged the beta is imperfect.

Suno introduced Speech in beta on October 1, 2026, in its official post titled "Introducing Speech (beta)." Suno The company already offers a separate Voices feature that lets users add their own voice to Suno-generated songs. That system checks a user's speech against an uploaded vocal recording and is not available to users under 18.

The broader context here is a change in how Suno builds its models. Suno replaced its AI models with a new model trained on licensed music. TechCrunch It rolled out AI models trained in cooperation with major music industry players. The company acquired WavTool for its AI music editing tools, a browser-based editor that launched in 2023. Suno raised more than $400 million at a $5.4 billion valuation. Reuters

In my view, the combined voice-and-music output is the part to watch. Most current tools create narration and background music as separate tracks that are mixed later. Doing both in one model run, what engineers call an inference pass, makes early drafts faster for solo creators. It can also cause problems with balance between voice and music, clarity, and style match. The music toggle and controls for gender, style, and variety help. The practical test will be consistency across takes, pronunciation in long scripts, and how clean the isolated voice track stays when music is on. If Suno tightens those controls during beta, Speech could become a fast drafting tool for formats where choosing music used to take manual work.