This is the curated IVA Swift dataset extracted from GitHub.
It contains curated Swift files gathered with the purpose to train a code generation model.
The dataset consists of 383380 swift code files from GitHub totaling ~542MB of data.
The uncurated dataset was created from the public GitHub dataset on Google BiqQuery.