This synthetic text dataset contains 100k samples of triples (english, morse, category), heavily utilizing Claude Sonnet 4.6.
Morse data was generated first. It contains simulated plausible transmissions that may happen over radiotelegraphy, particularly amateur radio.
English interpretations (a.k.a. "translation") are generated from the Morse ground truth.
The dataset samples follow four categories:
Ragchew (70%): Casual conversations between… See the full description on the dataset page:
https://huggingface.co/datasets/michaellin/morse-translate-transcribe-100k.