On AI Music

(screenshot from video game F-Zero GX / Nintendo / Sega / Amusement Vision)

31 AUG 26

I’ve written previously that Creative Inspiration is the most fruitful force in the universe. We need to keep it that way.

The technology is here. As creators, we shouldn’t bury our head in the sand. Rather, we should find sustainable ways forwards that keeps Creative Inspiration thriving amongst us, and not let a few powerful AI companies centralize the means to “creative” technology. These companies shouldn’t reap all the benefits for their reckless and exploitative rush to market, damaging creators and creative economies through their lack of collaboration along the way. This technology needs to be back in the hands of the people.

I’ve been trying a number of different GenAI tools and workflows over the last few weeks to get a grounded perspective on where things are at. Whereas a month ago I was AI-Skeptical – thinking that it was unreliable and not worth the power consolidation and environmental costs – a few things have changed:

1) Open-Weight LLM models have now hit frontier level ability. This means that anyone can download and run frontier-level models freely. This alone will stop monopolies emerging from Silicon Valley, and levels the playing field worldwide.

2) While electricity is still a major concern, Data Centre water technologies have gotten much more efficient over the last year. It shows real work is being done for environmental concerns, and I believe this will continue.

3) I’ve experienced, first hand, how we can leverage these technologies to make human-empowering technology. I’ve written about this in my post, AI Building Exoskeletons, and one my current projects is using AI-Assisted programming to make it easier for listeners to understand how I compose my chiptune music. This is an example of how we can use AI to build freely available, educational, empowering, and reliable tools.

With the above context, I’ve now tested the current state of music-generation AI tools. Whereas I couldn’t see myself using many of the features in creative process, it is in the “Cover” function that my mind has been blown. I am simultaneously feeling fear for the countless ramifications, as well as creative inspiration for the possibilities it enables.

One of my favourite compositional workflows is composing 4-Channel Game Boy chiptune music. It is a scope that fits my personality and energy levels well, and it focuses that energy on melodic and harmonic composition. So to hear one of my chiptune compositions translated into an impressive sounding variety of genres (see above video) – let’s just say it has had a genuine impact on me.

Similar to what I’ve discovered with AI programming workflows, the key between unreliable and reliable AI is in starting with a good core idea, and in iterating on the outputs. Using AI is like using jet fuel: It will move you in a given direction much faster. The key is in making sure you’re heading in the right direction to start with. This means you can spend millions of tokens / energy / time creating the wrong things if you’re not being mindful. Where you start getting interesting and “creative” results out of AI is in the fact this speed lets us very quickly prototype, test, and iterate on ideas. You repeat this cycle until the results get closer to what you have in mind. Or even not in mind – it might simply be a process of exploration and discovering new possibilities we couldn’t previously audiate.

I think the single most impactful experience I’ve had like this is with the first track in the above video. I started with my own original chiptune composition, flight_school, and then passed it through a few iterations of different genres to get an interesting and uncommon blend of genre-specific-features that resonated. It is a process of asking “What If?” on steroids. Hearing the genre-bending cover of my own track has truly creatively inspired me, and has made me want to seek out ways of recreating it through a more transparent and shareable workflow.

Cases like the above is why I don’t think we should be against the technology itself. GenAI can open up truly transformative workflows, BUT, it cannot stay in the hands of the few companies who likely spent more money on lawyers than on paying artists (looking at you Suno and Udio). What NEEDS to come next is that this technology, which already exists, is transformed into Open, freely downloadable versions. This technology needs to be given back to the artists it stole work from. This has already happened in the LLM space (see DeepSeek, Kimi, Qwen, etc). It needs to happen in the artist space. People fear what might happen if everyone has access to this technology without limit. But at least, in that world, the young artist without venture capital money can own the technology, not sending data or money anywhere else, and adapt it to their own local working conditions and process. This is the path I see forwards.

Until truly Open music models arrive, I don’t think I’ll be using it in my production process. As of today, AceStep 1.5 has been the only significant advancement in this direction, but we are still not nearly frontier level. As a creative community, Open models are how we can actually be empowered by this technology.

In the meantime, I will continue exploring how we can leverage other GenAI technologies in ways that actually educate and empower creators, rather than replacing them. I understand if you don’t want to join me on this journey. But if you are also interested in how we can bend this technology back towards Creative Inspiration, keep in touch and watch this space!

Related:


Location / Album: