Pre-tokenized training data used to train soyrsoyr/erebus-v2-1.5b-base, a 1.5B parameter Qwen3-architecture language model trained from scratch.
Binary files containing packed token IDs as uint32 values (4 bytes per token), tokenized with Qwen/Qwen3-1.7B tokenizer (151K vocab). Each document is separated by an EOS token.
Files can be loaded as memory-mapped numpy arrays:
import numpy as np
tokens =… See the full description on the dataset page:
https://huggingface.co/datasets/soyrsoyr/erebus-v2-training-data.