This is a sampled subset of PleIAs/SYNTH containing approximately 109,149,965 tokens.
Original Dataset: PleIAs/SYNTH (~87B tokens, 79.6M samples)
Sampling Method: Reservoir sampling (unbiased random sampling)
Target Token Count: 100,000,000 tokens
Actual Token Count: 109,149,965 tokens
Tokenizer: GPT-2 (50,257 vocabulary)
Documents Sampled: 100,000… See the full description on the dataset page:
https://huggingface.co/datasets/codelion/synth-100M.