This repository provides a tokenized version of the English split of the openbmb/Ultra-FineWeb dataset, prepared for large-scale language model training. The dataset consists of 100 billion high-quality tokens, processed with a custom tokenizer.
The data is sharded into 100 files, each containing exactly 1 billion tokens, making it easy to stream and use in distributed training setups.