These subsets were produced by filtering the original Pick-a-Pic v2 dataset using FiFA, a data filtering algorithm proposed in the paper Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion Models.
Preference Margin: Estimated using PickScore
Text Quality: Estimated through LLM scoring
Text Diversity: Estimated using K-NN distance… See the full description on the dataset page:
https://huggingface.co/datasets/Dragonjinny/FiFA-pickapic-v2.