This is a sampled subset of PleIAs/SYNTH containing approximately 14,631,489 tokens.
Original Dataset: PleIAs/SYNTH (~87B tokens, 79.6M samples)
Sampling Method: Reservoir sampling (unbiased random sampling)
Target Token Count: 10,000,000 tokens
Actual Token Count: 14,631,489 tokens
Tokenizer: GPT-2 (50,257 vocabulary)
Documents Sampled: 13,345
Documents… See the full description on the dataset page:
https://huggingface.co/datasets/codelion/synth-10M.