Toro y Moi is on stage at Google I/O 2025, working a MIDI controller connected to Google’s Lyria RealTime model. No samples loaded. No pre-recorded loops queued up. He’s mixing text prompts like “chaotic conga player” and the model is generating music. Live. In real time. The sound changes as he adjusts the weights between prompts, and what’s coming out of the speakers has never existed before.

“Everything I’m playing is completely improvised,” he says. “I’m basically jamming with the computer. The computer will also be jamming with me.”

I’ve been building instruments and systems for live electronic music since the early 2000s. VST plugins, custom MIDI controllers, percussion instruments, granular samplers. All of them trying to answer the same question: what does it mean to perform electronic music live? Watching this pre-show, something clicked. This wasn’t playback or triggering. The music was being generated on the spot, from words.

Down the Rabbit Hole

I went from watching to building within hours.

Google had released PromptDJ MIDI in AI Studio, the same kind of interface Toro y Moi used on stage. Forkable. I forked it, started reading the API docs, and built my own version in a Google Colab.

The model streams 48kHz stereo audio with about two seconds of latency between a control change and hearing the result. You steer it with weighted text prompts. You can lock it to a key, set tempo, control note density, mute instrument groups. A full instrument hiding behind an API.

I tried locking it to a key and feeding it different prompts within that constraint. I tried slowly crossfading between two completely different descriptions. The thing that got me was the transition. You type a new prompt and the music doesn’t jump. It drifts. You hear the old prompt dissolving and the new one taking shape, note by note, texture by texture.

I played with this for days.

What I Made

Two ambient pieces I generated with the Colab. No editing, no post-processing. Prompt in, audio out.

Demo 1

Demo 2

A New Kind of Instrument

Every instrument I’ve built was some version of the same idea: give a person controls, map those controls to sound parameters, let them perform. A knob turns and a filter sweeps. A pad gets hit and a sample fires. Sometimes there’s randomization, sometimes there’s quantization to catch bad timing, but the sounds are always pre-made. The performance lives in the arrangement.

Lyria RealTime does something different. The sounds don’t exist until you describe them. You’re not triggering pre-loaded responses. You’re feeding the model language and it’s producing music it thinks fits. The performer’s job shifts from triggering and mixing to describing and steering.

That’s not a synthesizer, not a sampler, not a sequencer. That’s a different thing altogether.

Back in 2008 I built a set of MIDI devices that let non-musicians jam on a composed piece. The whole point was lowering the barrier: quantize the input, hide the complexity, let anyone play. Lyria takes that further than I thought was possible. You describe what you want to hear, and you hear it. No musical training required. No knobs to learn. Words in, music out.

Where This Goes

Live shows where the audience sends prompts that feed into the music. A performer on stage steering the model while the crowd shapes what comes out. That’s a gig I want to play.

Then there’s the duo angle. A musician and the model improvising together. You play something, the model hears it and responds. You respond to that. Closer to playing with another person than any backing track or loop station has ever been.

And at the other end of the scale: a box on your desk that generates ambient music all day. Not a playlist. Music that adapts to a prompt you set in the morning and drifts on its own. Hardware running a local model, no internet needed. That one’s further out. These models are still big and cloud-bound. But you can see where it’s heading.

Right now, though, sessions cap at 10 minutes. You’re in the middle of something good, the textures are layering, you’re finding a groove, and the connection drops. You can restart, but the context is gone. Like someone unplugging your guitar mid-solo and handing you a different one.

Ten minutes at a time. That’s the leash. I keep going back anyway.

music, ai, opinion