5816 model responses = 8 benchmark runs × the 727 stimuli of
Full-Duplex-Bench v1.0 (pause handling,
backchannel, smooth turn-taking, user interruption). The runs cover 4 systems;
several differ only in voice prompt, prompting regime or weights, which is the point — those are
controlled pairs. For every stimulus and run you get the
model's own reply channel as lossless FLAC — time-synchronous with the stimulus, so… See the full description on the dataset page:
https://huggingface.co/datasets/MagicLuke/fdb-v1-outputs-v1.