A synthetic spoken-dialogue DPO (Direct Preference Optimization) dataset with 8482 examples,
designed for preference alignment of textless speech language models.
Each example corresponds to a single assistant turn in a two-speaker conversation.
For every turn the dataset provides:
The spoken prompt (the last user utterance, as audio + text)
A chosen response: the best model-generated spoken reply (scored by an LLM… See the full description on the dataset page:
https://huggingface.co/datasets/vprak17/synthetic_audio_paired_preferences.