· via The Verge
Suno launches Speech beta, generating AI voiceovers with optional background music
Suno's new Speech feature, now in public beta on web and mobile, turns prompts or scripts into spoken audio, optionally paired with AI-generated background music in a single track.

Suno moves from music into speech
Suno, the company best known for its AI music generator, has released a feature that creates spoken audio rather than songs. According to The Verge, Speech is available in public beta now across Suno's web and mobile apps, and it can produce a voiceover and its background music together in a single generation.
The combination is the point, Suno argues. Chief product officer Jack Brody described it in the announcement as "the first audio model that generates voice and music together as one cohesive track," adding that music remains at the heart of the company even though its ambitions have always stretched to other kinds of human expression.
What the feature does
Speech sits under the Create tab and comes in two modes. Simple mode takes a plain-language description of what you want — The Verge gives the example of a pirate captain rallying his crew — while Advanced mode accepts a custom script when you already know the exact wording. Advanced settings add control over the gender of the generated voice, the speech style, and how much variety successive generations introduce. A single output can run up to roughly eight minutes.
The background music is optional rather than mandatory: a toggle removes it when only clean speech is needed. Suno frames the pairing as a fit for particular spoken formats, such as a quiet score beneath a poem or something more energetic behind dramatic voiceovers and encouraging speeches.
The company is candid about rough edges. "Beta really does mean beta," Brody said, admitting that a British accent can occasionally drift into an Australian one and that dramatic pauses may turn out to be very dramatic. Suno says it will keep improving Speech based on user feedback.
A familiar market with a twist
As The Verge points out, AI-generated speech is crowded ground rather than new territory. DeepMind has experimented with deep-learning speech synthesis for a decade, Adobe offers a text-to-speech tool, and ElevenLabs has become one of the most recognizable platforms in the category since launching in 2023. Suno's differentiator is not the speech itself but the fusion of voice and music generation inside one product.
There may also be a defensive logic to the move. The Verge suggests the expansion is likely an attempt to diversify the platform, given that Suno's music generator has attracted numerous lawsuits.
Why it matters
For creators, the appeal is the collapsed workflow: narration and a score usually come from separate tools and have to be lined up manually afterward, and Speech turns that into one generation step. If quality holds up as the feature matures beyond beta, it becomes a fast route to finished audio for narrated video, poetry, presentations and similar formats, while the music toggle keeps it viable for plain voiceover work.
For the industry, the release signals that Suno wants to be an audio company, not strictly a music company. That puts a prominent generative audio brand into the lane of speech specialists like ElevenLabs and Adobe, with a bundled voice-plus-music capability those tools do not currently match. And with the music generator entangled in litigation, a second product line gives Suno somewhere to grow regardless of how those legal disputes resolve.
- #suno
- #text-to-speech
- #generative-audio
- #ai-music
- #public-beta