Views
No views yet
HelpSteer2 with four objectives (helpfulness, correctness,
coherence, conciseness). Preferences come from a prompted pairwise oracle (Qwen3-32B, one
attribute rubric per comparison, swap-averaged over presentation order); no reward model is fitted.
Base and reference are meta-llama/Llama-3.1-8B-Instruct.Phi-4 and Llama-3.3-70B), neither of which supplied the training preferences:
worst objective 0.369, average 0.410. Single seed.