Speech dataset prepared with Trelis Studio.
Metric
Value
Source files
1
Validation samples
9
Total duration
3.3 minutes
Column
Type
Description
audio
Audio
Audio segment (16kHz) - speech only, silence stripped via VAD
text
string
Plain transcription (no timestamps) - backwards compatible
text_ts
string
Transcription WITH Whisper timestamp tokens (e.g., `<
start_time
string
Segment… See the full description on the dataset page:
https://huggingface.co/datasets/Trelis/latent-space-validation.