musegen
This is a small Transformer that continues melodies. It runs in this page: the 810,240 weights load into your browser and every note below is sampled here, in JavaScript. Drag the red numbers sideways (or tap them) to change the input.
Take the first 4 bars of held-out piece 1 and let the model write 12 more bars in the prompt's key. Sample at temperature 0.90 from the top 95% of the probability.
| Model | Real |
|---|
I started this in COGS 319 (Fall 2024) as an LSTM in a Colab notebook. Later I rebuilt it as a tested PyTorch package with three models, a better encoding of music and an honest evaluation. The model on this page is the package's Transformer: 4 layers, 128 dimensions, trained for 25 epochs on a CPU. Its perplexity on held-out pieces is 3.00.
Be clear about the data. The training set is 256 short pieces that the package composes itself (sources.py): each has a key, a chord progression and a repeated phrase. That keeps the demo free of copyrighted music, and it means the model writes in that simple style. The 26 prompts above are the test split, which the model never saw.
What the model reads
The model never sees audio or a piano roll. Each melody becomes a list of events, following REMI
(Huang and Yang, 2020). BAR starts a bar. POS_n puts the next note on the
nth sixteenth of the bar. VEL, PITCH and DUR then give its
loudness, pitch and length. These are the tokens of the melody above; the prompt is grey. Click a
red token to see how the model picked it.
This turns composing into choosing the next token out of 148, the same task a text model solves. It also means the model can write nonsense: a pitch with no length, or a position earlier than the last one, which does not decode into music. The encoding lives in tokenizer.py.
Picking one token
At each step the model gives a probability for every token. Before drawing one, I filter that distribution three times. A grammar keeps only the tokens that can legally come next, so every sequence decodes. A key constraint keeps only pitches in the scale. Top-p sampling keeps the smallest set of likely tokens whose probabilities add up to p (Holtzman et al., 2020). Below is the full distribution at one step of the melody above. Faint bars are tokens the grammar rules out; grey ones are pitches outside the key.
. Context:
| Token | Model | After filters |
|---|
Raise the temperature and watch that worst step grow. A strange draw, like a note 23 sixteenths long, confuses the model, and it then asks for a second length in a row or a position it has already passed. The mask is what keeps those steps from breaking the melody.
The mask is cheap because the grammar is regular. A small state machine tracks where the sequence
is (after a bar line, inside a note, at which position) and returns the set of legal next tokens
(GrammarState in
tokenizer.py).
The sampling loop is generate_tokens in
generation.py.
While building this page I found a bug there. If a
prompt ended in a rest, the grammar still allowed positions later in its last bar, and the model
filled them, writing notes into the prompt. With the 26 test prompts cut two beats early, all 52
continuations did. Now the sampler opens a new bar first and none do. A test checks it.
Finding the key
The constraint needs a key, and the prompt does not say what it is. I estimate it with the
Krumhansl-Schmuckler algorithm. Add up how long each of the 12 pitch classes sounds. Then correlate
those 12 numbers with the probe tone profiles that Krumhansl and Kessler (1982) measured by asking
listeners how well each note fits a key, rotated to all 24 major and minor keys. The best
correlation wins. The code is estimate_key in
theory.py.
Does the constraint help?
This runs the model on all 26 held-out prompts twice, once constrained to the key of the prompt and once free, and compares both with what the pieces really do next. Each prompt gets the same seed both times.
| Notes in scale | Key clarity | Mean leap | Repeated figures |
|---|
"Notes in scale" is the share of notes inside whichever major or minor scale fits that melody best. "Key clarity" is its best correlation from the chart above. "Repeated figures" counts 4-note patterns of intervals and lengths that occur more than once, so a motif restated higher still counts. These follow the set comparisons of Yang and Lerch (2020) and are computed by metrics.py in the package and its port in this page.
The constraint fixes the wrong notes without retraining. It does nothing for structure. The real pieces restate their phrases and the model almost never does. That is the clearest gap, and the one I would work on next, with more data and a longer training context.
The rest of the pipeline
Reading MIDI. The notebook called get_piano_roll(fs=4) to get
"4 steps per beat", but fs is frames per second, so the grid changed with the tempo.
The package quantizes against each file's own beat grid, keeps the top note at each onset as the
melody, ignores drum tracks and repairs files with broken key signatures
(midi_io.py).
The model. A decoder-only Transformer with pre-LayerNorm blocks, rotary position embeddings (Su et al., 2021) and a key/value cache, so each new token costs one step instead of a pass over the whole context (transformer.py). For this page I exported the weights as float16 and rewrote the forward pass in plain JavaScript (model.js). A test runs both on the same tokens and checks that the logits agree (test_web_parity.py). The same test covers the tokenizer, grammar, sampling filter, key finding and metrics.
Evaluation. The notebook shuffled overlapping windows and then split them, so validation windows shared 47 of 48 frames with training ones. The package splits by file. The README has the full benchmark against two LSTM baselines and a list of the fifteen problems I fixed from the original notebook.
Run it yourself
git clone https://github.com/edwin580/music-generator.git
cd music-generator
pip install -e ".[dev]"
musegen synth-data --dest data/synthetic --n 256
musegen train -c transformer --midi-dir data/synthetic
musegen generate -m runs/transformer/best.pt -o song.mid --key auto --wav
Or train on a free GPU with the Colab notebook.
References
- Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. (2020). The curious case of neural text degeneration. ICLR.
- Huang, Y.-S., and Yang, Y.-H. (2020). Pop Music Transformer: Beat-based modeling and generation of expressive pop piano compositions. ACM Multimedia.
- Krumhansl, C. L., and Kessler, E. J. (1982). Tracing the dynamic changes in perceived tonal organization in a spatial representation of musical keys. Psychological Review, 89(4), 334-368.
- Su, J., et al. (2021). RoFormer: Enhanced Transformer with rotary position embedding. arXiv:2104.09864.
- Yang, L.-C., and Lerch, A. (2020). On the evaluation of generative models in music. Neural Computing and Applications, 32, 4773-4784.
- Victor, B. (2011). Explorable Explanations. The draggable numbers on this page follow his Tangle library.