Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Swift, a popular statically typed language.
Target Language: Swift
Dataset Size:
Selected Swift as the target language due to its… See the full description on the dataset page:
https://huggingface.co/datasets/ammarnasr/the-stack-swift-clean.