Consumer software spent years treating music as something people played back rather than something they drafted in a chat box. That changed as text-to-audio models moved out of research papers and into everyday apps.
Streaming services now report a large share of new uploads as fully generated, and the tools themselves have shifted from thirty-second loops toward songs with verses, choruses, and a stated length.
The results still sit in a legal and cultural grey zone, but the product surface is no longer experimental.
Google DeepMind's work in this area ran through MusicLM and then the Lyria family, which first showed up in tools such as MusicFX and YouTube's Dream Track. Lyria 2, later listed as lyria-002 on Vertex AI, produced short clips and leaned instrumental. Lyria 3 reached the Gemini app in February 2026 with vocals, auto-written lyrics, and image prompts, first at about thirty seconds. A Pro variant in March extended that into structured tracks of roughly three minutes.
Then, Google announced 'Lyria 3.5.'
Lyria 3.5 first arrived in July inside Google Flow Music.
On 4 September 2026, Google opened the same model in the Gemini app and the Gemini API, on web and mobile, with additional access through Google AI Studio and Google Vids.
Lyria 3.5 is built for full-length songs rather than hooks.
Google describes gains in three places: musicality, lyrics, and vocals. Arrangements are meant to hold more complex melodic structure and a more natural movement from section to section.
Lyrics are generated with tighter prompt following and an awareness of song form, so a track can include verses, choruses, and bridges instead of a single repeating phrase.
Vocals are described as more expressive and emotionally varied, with clearer pronunciation.
Users can ask for a sung track or an instrumental, name or describe a genre, and set tempo.
Duration is variable: a one-minute clip, a mid-length piece, or a cohesive song up to three minutes.
In Gemini, templates cover everyday jobs such as background beds and birthday songs. Prompts can stay loose (mood, memory, style) or get specific (BPM, instruments, vocal timbre).
The model also accepts text and images.
In the API, a prompt can be paired with up to ten images, and the audio is composed against that visual material. Developers can supply their own lyrics with section tags and use timestamps to mark when instruments enter.
Output is 44.1 kHz stereo. MP3 is the default; WAV is available for Lyria 3.5.
Lyrics and vocals can follow the language of the prompt.
A separate Clip model in the same family still exists for fixed thirty-second pieces. Generation is single-turn: once a track is produced, it is not edited across follow-up prompts in the same way a chat model revises a paragraph.
And just like other AI tools, results generated by Lyria 3.5 carry an embedded SynthID watermark.
The watermark is embedded in the audio so the clip can be flagged as machine-made after compression or light editing. Among others, this watermark can trigger filters to block requests from people who ask for a named artist's voice or for copyrighted lyrics.
If a prompt names an artist, the system is instructed to treat that as broad stylistic reference rather than a clone. Those limits sit against a wider dispute over training data and liability, including a Munich court judgment against Suno that is not final and that Google has not tied to this release.
In practice the model is aimed at backing tracks, short branded cues, ringtones, and first-pass songs that can be exported into other tools.
Flow Music remains the more production-oriented surface, with extra controls such as BPM and stem export on longer tracks. Gemini is the simpler path: pick a length and a vocal setting, write a prompt or drop in a photo, and take the file that comes back. How close any given output sits to a finished record still depends on the prompt and on what the listener is willing to accept as a song.





















































































































































































































































































































































































