This repository contains a pre-downloaded and processed version of the Stack-Edu dataset. The original dataset consists of SWHIDs pointing to files on Software Heritage. This version contains the full text content, saved in Parquet format and organized by programming language, making it immediately usable for training and analysis without the slow download process.
Stack-Edu is a 125B token dataset of high-quality, educational code filtered… See the full description on the dataset page:
https://huggingface.co/datasets/meryyllebr543/stack-edu-huggingface.