Suno Introduces 'Speech' Feature Combining Spoken Audio and Background Music
AI music startup Suno has launched a new feature called Speech, allowing users to generate spoken text alongside matching background tracks in a single audio output.

AI music generator Suno is expanding beyond pure song generation with the release of a new feature called "Speech." The tool produces spoken-word audio and matching background music together, outputting both elements as a single integrated audio track.
Generating Voice and Soundtracks Together
To use the feature, users type in text or an idea and describe their preferred voice and musical style. The model then generates both the voiceover and the complementary background sound.
According to Suno product chief Jack Brody, the company evaluated Speech with a small test group for a month prior to the broader release. Suno suggests that users can leverage the capability to produce content such as poems, guided meditations, and bedtime stories.
Because the feature remains in beta, Suno noted that it still experiences occasional bugs, including voice inconsistencies where a British accent may intermittently sound Australian.
Training Questions and Ongoing Legal Pressures
Suno has not disclosed details regarding how it trained the underlying model powering the Speech feature.
The update comes as generative music platforms face mounting scrutiny over the sources of their training data. Major record labels have filed lawsuits against Suno over potential copyright infringement. Additionally, a Munich court recently ruled against the startup, rejecting fair use arguments as a legal justification for utilizing copyrighted material.



