The complete open dataset behind MacWispr's on-device
dictation polish model (Qwen3.5-0.8B post-trained to turn raw speech-to-text into
clean, structured writing). Training pipeline and verifier live in the
MacWispr repo.
sft/train.jsonl (+valid/test)
3,011 / 276 / 173
Main SFT pool. {"text": "### Input:\n\n\n### Output:\n"}
synthetic/synth_hard.jsonl
311
Synthetic… See the full description on the dataset page:
https://huggingface.co/datasets/vasanth009/macwispr-polish-data.