A translation of HuggingFaceH4/ultrachat_200k (train_sft split) into the Deseret Alphabet — a 19th-century phonetic writing system for English, encoded in Unicode at U+10400–U+1044F.
This dataset is intended for supervised fine-tuning (SFT) of chat models that operate exclusively in the Deseret Alphabet.
Conversations: 200,000 multi-turn user/assistant exchanges
Format: JSONL, one conversation per line… See the full description on the dataset page:
https://huggingface.co/datasets/chrisjpatty/ultrachat-deseret.