🌐 DemoPage | 🤗SFT Dataset | 🤗 Benchmark | 📖 arXiv | 💻 Code | 🤖 Chat Model | 🤖 Base Model
MusicPile is the first pretraining corpus for developing musical abilities in large language models.
It has 5.17M samples and approximately 4.16B tokens, including web-crawled corpora, encyclopedias, music books, youtube music captions, musical pieces in abc notation, math content, and code.
You can easily load it:from datasets import load_dataset
ds =… See the full description on the dataset page:
https://huggingface.co/datasets/m-a-p/MusicPile.