This dataset converts the Lakh MIDI Dataset into a structured text format inspired by the Multitrack Music Machine (MMM) paper. It includes 344,900 samples, each representing an 8-bar symbolic music fragment, tokenized into a language-model-friendly format.
Each line in the dataset is a music fragment composed of tokens like:
PIECE_START… See the full description on the dataset page:
https://huggingface.co/datasets/juancopi81/lmd_clean_8bars_32th_resolution.