This is a filtered version of bigcode/the-stack-dedup that subsets for the following language splits:
[
'ada', 'assembly', 'awk', 'c', 'cpp', 'c-sharp', 'clojure', 'common-lisp', 'dart',
'erlang', 'f-sharp', 'fortran', 'glsl', 'go', 'haskell', 'isabelle', 'java', 'javascript',
'julia', 'kotlin', 'llvm', 'markdown', 'mathematica', 'pascal', 'perl', 'php',
'powershell', 'python', 'r', 'racket', 'ruby', 'rust', 'scala','shell', 'solidity',
'sparql', 'sql', 'stan', 'typescript', 'zig'
]… See the full description on the dataset page:
https://huggingface.co/datasets/MaLA-LM/stack-final.