A proposal-writing preference dataset (DPO) derived from the
CORDIS database of EU-funded research projects.
It trains a model to prefer high-quality, funded-style proposal text over degraded
alternatives, refining writing quality after supervised fine-tuning.
Conversational preference format compatible with TRL DPOTrainer:
{
"prompt": [{"role": "user", "content": "Draft the 'Objectives and Scope' section… See the full description on the dataset page:
https://huggingface.co/datasets/RCaz/eu-funding-proposals-dpo.