Pairwise human preference labels for AI-generated educational video tutorials.
Source: Expert annotations from creators.metaphi.ai Video Arena
Format: Each row is one A/B comparison per chapter with winner label and reasoning
Agents: Three harnesses (Claude Code, Codex, Gemini CLI) compared pairwise
Use case: RLHF reward model training, preference-based optimization