Strictly parsed and reserialized JSON pretraining data curated from the json
subset of bigcode/starcoderdata at revision 9fc30b578cedaec69e47302df72cf00feed7c8c4.
Tokens: 500,000,142 with HuggingFaceTB/SmolLM2-135M
Rows: 1,115,107
Binary: headerless little-endian uint16, EOS-delimited
Validation: see validation_report.json
Samples: see samples/high_quality_samples.jsonl and .md
This is a derivative of StarCoderData and remains subject to its… See the full description on the dataset page:
https://huggingface.co/datasets/glouriousgautam/lilm1-structured-json-500m.