A curated dataset for training speech-to-text cleanup models to achieve optimal transcript refinement.
This dataset contains paired examples of raw speech-to-text transcriptions and manually-cleaned versions, designed for fine-tuning models to clean up transcripts to a specific quality level ("Goldilocks" cleanup - not too much, not too little).
dataset/
├── data/
│ ├── audio/… See the full description on the dataset page:
https://huggingface.co/datasets/danielrosehill/Transcription-Cleanup-Trainer.