Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Selected Rust as the target language due to its… See the full description on the dataset page:
https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.