This dataset contains expert votes collected in the text-only category. Each row represents a single vote judging two models (model_a and model_b) on a user conversation, along with the full conversation history. Key fields include:
id: Unique feedback ID of each vote/row.
evaluation_order: Evaluation order of the current vote.
winner: Battle result containing either model_a, model_b, tie, or both_bad.
conversation_a/conversation_b: Full conversation of the current… See the full description on the dataset page:
https://huggingface.co/datasets/sanderland/arena-expert-5k.