Continued-pretraining (CPT) corpus in a constructed language generated by
the ConlangCrafter pipeline (language id
bd412d52).
This dataset exists to give language-model fine-tuning runs a demonstrably
out-of-distribution target. Because the language was synthesized by an LLM
pipeline after public model pretraining cutoffs and never published in any
form before, no large pretrained model has seen it. CPT against this… See the full description on the dataset page:
https://huggingface.co/datasets/TearedModels/conlangcrafter-cpt-bd412d52-hard-typo2.