Phase 2 of 3: Higher-quality filtered data with context extension (250B tokens) used for mid-training of Ettin models.
This dataset contains the mid-training phase data used to train all Ettin encoder and decoder models. This phase focuses on higher-quality filtered data and context length extension to 8K tokens. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.
📊 Data Composition… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-extension-data.