High-quality educational text data for language model pretraining, derived from
the OpenGloss
synthetic encyclopedic dictionary and related curriculum materials.
This dataset is derived from OpenGloss, a synthetic encyclopedic dictionary
and semantic knowledge graph for English that integrates lexicographic definitions,
encyclopedic context, etymological histories, and semantic relationships in a
unified resource. OpenGloss… See the full description on the dataset page:
https://huggingface.co/datasets/mjbommar/oglm-curriculum-pretrain.