This dataset consists of 100k audio preference pairs generated by TangoFlux during the CRPO stage. Specifically, TangoFlux performed five iterations of CRPO. In each iteration, 20k prompts were sampled from a prompt bank. For each prompt, audio samples with the highest and lowest CLAP scores were selected to form the "chosen" and "rejected" pairs, respectively. This process resulted in a total of 100k preference pairs.
Since every iteration contains 20k prompts… See the full description on the dataset page:
https://huggingface.co/datasets/declare-lab/CRPO.