Lyria
Google DeepMind's AI music generation model — produces high-fidelity instrumental audio with fine-grained creative control.
3 min read · updated Aug 2026
💡 In plain words
An AI music model from Google that creates professional-quality instrumental music from text descriptions. You describe the genre, instruments, and mood, and Lyria generates a polished audio track. Built for developers and musicians who want fine control over what the AI produces.
🎯 A real example
Need a cinematic piano piece with strings? Describe it to Lyria with your desired tempo and mood, and get a high-fidelity instrumental track generated by Google's AI.
🤔 Is it for you?
- Developers building music features into apps via API
- Musicians who want fine control over instruments and tempo
- Projects that need Google-ecosystem integration
- You need vocal generation with lyrics
- You want a free, local, open-source solution
- You prefer a simple web interface over API integration
What is Lyria?
Lyria is Google DeepMind’s AI music generation model family, designed to produce high-fidelity instrumental audio from text descriptions. First announced in April 2025 as Lyria 2, the model has since evolved through versions 3 and 3.5, each adding improved quality, longer composition capability, and finer creative control.
Unlike consumer-facing generators such as Suno or Udio, Lyria is a developer and enterprise tool — available through Google’s Gemini API rather than a simple web interface.
How Lyria works
You describe what you want — genre, mood, instruments, tempo — and Lyria creates a track that sounds coherent from start to finish. Themes stay consistent, harmonies progress naturally, and rhythms hold together across the full piece.
What sets Lyria apart is how much control you get through the API. You can set exact BPM values, pick specific instruments, choose key signatures, and define song structure. This makes it a great fit for apps that need tracks matching exact specs — background scores for video, game soundtracks that shift with gameplay, or branded content with a specific sonic identity.
Model versions
Lyria 2 was the initial public release in April 2025, establishing the model’s foundation for high-fidelity instrumental generation through the Gemini API.
Lyria 3 expanded the model’s capabilities with improved tonal consistency and support for more complex arrangements, including multi-instrument compositions that maintain coherent interplay between parts.
Lyria 3.5 is the current flagship, offering the longest compositions and highest fidelity. It supports extended form generation — pieces that develop thematically over several minutes rather than simply looping or drifting.
Each version has been available through the same API endpoints, so developers can upgrade models without changing their integration code.
Key features
Fine-grained creative control over instruments, tempo (BPM), genre, and overall feel — more precise than most text-to-music tools offer through their interfaces.
Professional-grade output suitable for real production work. The Large model targets quality on par with studio-produced compositions.
Safety measures including content filters, recitation checking, and artist intent checks. These prevent the model from creating content that too closely copies existing copyrighted works.
SynthID watermarking on all generated tracks. This invisible mark can be detected by software, helping platforms and users spot AI-created content.
Google ecosystem integration through the Gemini API, making it straightforward for developers already using Google Cloud services to add audio generation to their applications.
Limitations
Lyria only does instrumental work — no vocals or lyrics. For full songs with singing, ACE-Step UI, Suno, or Udio are the right tools. It’s not open source, so you can’t run it locally like ACE-Step 1.5 XL or Stable Audio’s open-weight models. And the API-first design means casual users will find consumer tools easier to use.
How Lyria compares to alternatives
Against Stable Audio, Lyria offers deeper Google ecosystem integration but lacks open weights for local use. Against ACE-Step UI, Lyria provides higher instrumental fidelity but no vocals, no local option, and no free tier. Against Riffusion, Lyria is vastly more capable but requires API access and payment.
For developers who want a node-based workflow that can incorporate Lyria alongside other AI models, ComfyUI offers an extensible visual pipeline environment, though Lyria integration requires custom API nodes.
Who is Lyria for?
Lyria is best for developers integrating AI audio into applications, enterprises building music features at scale, and musicians who want precise instrumental generation through an API. For full songs with vocals, ACE-Step UI or Suno are better choices. For a free, local alternative, ACE-Step UI is the clear winner.
Conclusion
Lyria represents Google’s serious investment in AI audio generation. While it’s more of a developer tool than a consumer product, its quality, parametric control, and seamless Google Cloud integration make it a powerful option for building audio features into applications and services at scale.
Get the best new tools — before everyone else
One short, friendly email whenever we add a tool worth your time. No spam, unsubscribe anytime.