A continuously growing DPO (Direct Preference Optimization) preference dataset spanning 10 knowledge domains. Generated automatically using the HuggingFace Serverless Inference API and uploaded incrementally — the dataset grows larger every day.
chosen: generated with a thorough, expert-level system prompt — detailed, structured, with examples and… See the full description on the dataset page:
https://huggingface.co/datasets/iPwnds/Project-Silkworm.