Raw eval outputs for the MNLP report Personality by Preference: Big-Five Control
in Small Language Models via Preference-Optimized Mixture-of-LoRA (Team Liberte).
Systems: prompted Base / Instruct references, P-React (persona-keyed
routing), DPO-Full, DPO-MoLoRA. Trained systems are 3 seeds (42/1/2).
judge/ — LLM-as-a-judge over the BFI, scored by GPT-4o-mini and Gemini-3.1-flash-lite
(N=25… See the full description on the dataset page:
https://huggingface.co/datasets/LiberteEPFL/lfm-eval-results.