
What happened
Suno launched Speech in public beta across its web and mobile platforms, letting users generate voiceovers and background music together from scripts or prompts. Chief product officer Jack Brody called it "the first audio model that generates voice and music together as one cohesive track."
Why it matters
Speech is Suno's bid to diversify beyond its music generator, which has attracted so many lawsuits, and pairs generated voices with music as one track rather than as separate text-to-speech output. Users can toggle the background music off for clean speech.
What to watch
Suno says "beta really does mean beta" and will keep improving Speech around user feedback, with a maximum duration of around eight minutes. Brody notes British accents can wander off to Australia and back, and dramatic pauses may be very dramatic.
WHO IT HITSCreators who already use generative audio for poems, voiceovers, dramatic readings, or speeches get a single tool that produces voice and soundtrack together, though the beta's travel-prone accents and eccentric pacing mean producers of polished commercial audio may still need to review output. The music industry and rights holders pursuing cases against Suno's music generator may also watch whether this diversification shifts the platform's legal exposure.
Summaries like this, in your inbox every morning.
Suno built its reputation on AI music generation, and that business has drawn so many lawsuits that the company now has reason to widen its surface area. Speech is that widening: it moves Suno from songs into spoken audio, a category that already includes DeepMind's decade of deep-learning speech synthesis experiments, Adobe's text-to-speech tool, and ElevenLabs, which launched in 2023 and has become one of the most recognizable platforms in the space. Suno's angle is not novelty in speech itself but packaging: voice and music generated together as one cohesive track, with a toggle for clean speech and controls over voice gender, speech style, and variety.
The company is upfront that the feature is far from perfect. Brody's own framing—"Beta really does mean beta"—concedes that British accents can wander off to Australia and back and that dramatic pauses may be very dramatic. Those are the sort of rough edges that matter most for use cases like poems with a calming soundtrack or energetic voiceovers and encouraging speeches, where tone is the product.
Whether Speech becomes more than a side feature likely hinges on how quickly Suno improves it around user feedback, as the company says it will, and on whether voice-and-music pairing proves genuinely useful rather than merely convenient. Creators experimenting with spoken-word audio and the rights holders already litigating against Suno's music side both have reason to watch how the beta develops.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Meta will link Naver Map's walking-navigation feature to its Ray-Ban Meta and Oakley Meta AI glasses in 2026…

Anthropic's new study says that if robot prices decline along past trends, it will take 40 years for robots to…

Flow Engineering announced a $50 million Series B at a $750 million valuation, co-led by Antonio Gracias and G…

Ramp economist Ara Kharazian says US firms are spending less on AI even as usage rose about 50 percent from Ju…

Microsoft AI launched MAI-Transcribe-2-Streaming, which Microsoft says ranks first for accuracy on Artificial…

ELYZA said it is launching "ELYZA RSI Research", and that for LLMs of 100 billion parameters or fewer it has r…
