A subset of the-stack dataset, from 87 programming languages, and 295 extensions.
Each language is in a separate folder under data/ and contains folders of its extensions. We select samples from 20,000 random files of the original dataset, and keep a
maximum of 1,000 files per extension.
Check this space for inspecting this dataset.
The dataset contains 87 programming languages:
'ada', 'agda', 'alloy', 'antlr', 'applescript'… See the full description on the dataset page:
https://huggingface.co/datasets/bigcode/the-stack-inspection-data.