One row per scored challenger turn in the Albedo subnet's published eval traces. Scoring is BINARY yes/no-question scoring: an evaluator writes a flat set of yes/no questions per task, and each judge answers them with 1/0 for the king and the challenger independently. A judge's yes_rate is the mean of its 1/0 answers; a side's score is the mean of its per-judge yes-rates. Every per-judge, per-side record that scored a turn is folded into the row's… See the full description on the dataset page:
https://huggingface.co/datasets/lumetix-ai/itorgov-sn97-albedo-eval-traces-v5.