SLM Synthetic DPO — short_factual_stop_behavior
Summary
Synthetic preference-pair dataset containing chosen and rejected assistant responses for short_factual_stop_behavior.
Dataset type: preference optimization
Total pairs: 6428
Signal: short_factual_stop_behavior
Language: English
Category: controlled_verbosity
Difficulty: 1
Template family: short_factual_answer
Eval family: short_factual_stop_behavior… See the full description on the dataset page:
https://huggingface.co/datasets/tohio/slm-synthetic-dpo-short-factual-stop-behavior.