End-of-turn (semantic VAD) turns built from word-level forced alignments, schema-compatible
with livekit/eot-bench-data.
Each row is one user turn: an audio clip (16 kHz mp3), its words, and ordered
silence_spans. Per the eot-bench convention the last silence span is the true
end-of-turn (eot); earlier spans are mid-turn hold pauses (labels positional, not stored).
For every data type, all shards except the last form the train base; that… See the full description on the dataset page:
https://huggingface.co/datasets/Scicom-intl/semantic-vad-eot.