sotto.app ·
Trained Model (bf16) ·
MLX 5-bit Model
Overview
124K+ synthetic training pairs for fine-tuning small language models on speech-to-text transcript cleanup. This dataset was used to train the SottoASR transcript cleanup model — a 350M parameter model that exceeds a prompted 2B model on this task while being 8x faster.
Part of SottoASR — a local, privacy-first speech-to-text application for macOS.
Task… See the full description on the dataset page: https://huggingface.co/datasets/juanquivilla/sotto-transcript-cleanup.