This repository contains the continued-pretraining data used for the OpenAster1 60B-token stage.
The data is stored in Megatron-LM indexed dataset format (.bin / .idx) and is organized into 1B-token shards. The original recipe targets 60B tokens. The audited trainable total is 54,940,010,598 tokens because the final shard is an accepted short shard.
shards/shard_000 to shards/shard_054: training shards.
validation/bin: channel-level… See the full description on the dataset page:
https://huggingface.co/datasets/binichallein/OpenAster1-data-60B.