This is the curated train split of IVA Kotlin dataset extracted from GitHub.
It contains curated Kotlin files gathered with the purpose to train a code generation model.
The dataset consists of 383380 Kotlin code files from GitHub.
Here is the unsliced curated dataset and here is the raw dataset.