The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device.
The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
I see so much in common with this project and the numerous AI-based UX design tools out there. Whether it's music or UI, now that the "generation" portion of the work costs zero, all that remains is taste.
And so much of taste comes from exploring and killing off possibilities that turn out to be dead-ends. I love the idea that models like these will help us find the dead ends faster, or even produce a gem here and there.
P.S. if you want another uncanny version of Fur Elise, listen to Beethoven's own 1822 revision: https://www.youtube.com/watch?v=s24TtiGgb6k. His 1810 version that we all know was simpler and more balanced. But for what it's worth, Beethoven didn't publish either of them.
One think I didn't see mentioned in the post- maybe I missed it- how large was the data? How many samples did you use to pretrain and post-train
I first parsed the "on device" as _on the piano_ and was intrigued how that worked :) maybe someone can make "dynamic" rolls for a player piano somehow (like one of those braille "screens")?
I think one of the great, early, joys of learning a piano is gaining the following intuitions: The seemingly harder path of learning sheet is actually faster. Your mind _should_ learn to think in two dimensions Spatial — where fingers go — and Time — pitch and tempo – ** when learning. The _internalization_ of Space and Time queues guide the fingers in a dance that is vastly satisfying. This skill leads you to the final part of the journey that is improvisation and the one more exciting than what i am on now.
---
**
Space: Your finger placement on keys right, e.g. knowing how to go from landmark/anchor notes(mid-c, G, F etc) and then go to the others above and below it. Crudely this is some what like typing from your landmark f and j qwerty keyboard
Time: The out singing/verbalizing of the notes/beats on a time measure as you play them(per the time measure). e.g. you can say out loud 1-2-3-4 for 4/4 measure, if the measure has quarter notes say out loud. And `1-e-and-a-2-e-and-a-3-e-and-a-4-e-and-a` for a 4/4 with 1/16th note granularity. Do this as you play the notes and you get a sense of tempo.
you could always do the discipline you liked before for self fulfillment purposes, and you can still do it for self fulfillment purposes now
unless they're actually serfs again and won't have bills to pay, but no access to anything outside of the fiefdom that they maintain all day
If you like such creative AI work, I would recommend looking at the other "NeurIPS Creative AI Track 2025" submissions as well.
Some protocols (like I2C or MIDI) are gonna be with humanity forever, probably.
The closest we've had to realtime orchestration around a melody in the "real world" is probably arranger keyboards though your left hand is still responsible for the chord progression itself.
[1] - https://en.wikipedia.org/wiki/Microsoft_Research_Songsmith
Would be fun to get a midi clock going and play some chords on my piano and have my synth start jamming along with the bass and my keyboard doing some performance. Or any combination of the above.
Does anybody know of a project that has produced more convincing results?
One feature request:
Instead of playing the AI-generated audio solely through the iPhone's speakers, add an option to send the audio as midi notes to a device (probably the same one you received the mini notes from).
Also consider checking their decisions about representation, etc
I can probably squeeze out quite a bit more than 100 notes/sec as well. I haven’t spent much time optimizing inference yet.
For the current model I’m using Core ML, which optimizes the kernels the first time you run it. I haven’t actually spent that much time tuning performance beyond that.
I guess no one actually wants to learn about harmony, about voicing, and voice leading, spend the hours. No one wants to learn how to actually play an instrument. I guess no one is willing to do the work.
They want to just press some keys and declare that they made what the computer generated.
This is all so depressing for me.
I think in the near future, the world will divide, and all of the people who wish to think and do, will put up a wall separating themselves from the hu-bots that infest the rest of the planet. It's alright I guess. I hope there isn't a slaughter of one side or the other.
I don't understand that as a takeaway. Even if this worked perfectly, it would not be meaningfully stopping or discouraging anyone from learning music, and no one is claiming that what this models outputs is something the player played. It's a toy that someone made presumably because they like piano.
But, but… wouldn't that be… (gasp) DISTILLATION?
Fun project!
You know, I was nodding along until you shoehorned that one in.
Also, yes, the orange fascist who attempted to coup his way to power, raped women, and is destroying democratic institutions is indeed bad.
Oh no, AI will take our jobs??? FUCKING LET IT!
Who even wants to do all these jobs if we don't HAVE to??
Ask politicians to give us UBI.
You want to attack the shit that could make shit easier instead of attacking the 200 year old institutions in place that ensure class divisions and perpetual debt and wage slavery? Smart buggers
I have bad news for you if you think that will improve with the current deployment of AI. You will also notice I didn’t mention anything about jobs
thing go plink
machine hear plink
machine make many more plink
man happy for plink is fun
man not hit machine with club or scream on orange site
man leave cave and touch plant
For anyone interested, I’d highly recommend reading Robert Gjerdingen’s article Gebrauchs-Formulas. https://www.researchgate.net/publication/259731561_Gebrauchs...
You can also listen to the transcript of four Russian composers, including Rachmaninoff, playing this pattern recognition and generation game at a dinner party in the late 1800’s: https://youtu.be/PlFPOWuwBHI?is=EKBK7QQkJs4MsTCU
Composers at the time could do this just by looking at sheet music and audiating, without using a piano.
There's a great story around how Beethoven, perhaps one of the strongest improvisers of his day, completely upstaged Steibelt, a contemporary musician and by all accounts a bit of a charlatan. I'll include the entire quote verbatim from "The Lives of the Great Pianists" by Harold Schonberg.