text-generation
text-classification
task_ids:
language-modeling
tags:
rlhf
dpo
preference-learning
ai-ethics
ai-safety
alignment
human-feedback
annotation
language:
en
size_categories:
n<1K
pretty_name: AI Ethics Preference Annotation Dataset
A human-annotated preference dataset for RLHF and Direct Preference Optimization (DPO), focused on AI ethics failure modes. 95 prompts, 190 response pairs, full… See the full description on the dataset page:
https://huggingface.co/datasets/philosophyFire/Ai_ethics_dataset.