We present here SpokenSwag as described in the paper "Slamming: Training a Speech Language Model on One GPU in a Day".
This dataset is based on allenai/swag and synthetised with 4 speakers from hexgrad/Kokoro-82M.
We show that perfoming DPO over the dataset can really improve performance of Speech Language Models.
We encourage you to also see the following resources, for further information:
Project Page:
https://pages.cs.huji.ac.il/adiyoss-lab/slamming/ Paper:… See the full description on the dataset page:
https://huggingface.co/datasets/slprl/SpokenSwag.