A preference alignment dataset of instruction/chosen/rejected triples, built for
Direct Preference Optimization (DPO) and RLHF-style alignment toward an academic,
authoritative Machine Learning tone.
Part of the LLM-ArXiv-Domain-Expert pipeline.
Built on the same extraction pipeline as ml-arxiv-instruct —
ArXiv papers parsed with docling, chunked into extracts. For each extract:
instruction: a… See the full description on the dataset page:
https://huggingface.co/datasets/danivpv/ml-arxiv-dpo.