Dataset Card for Code Corpus code-v1
Dataset Description
Dataset Summary
Code Corpus code-v1 is a token-balanced collection of source code, commit messages and
diffs, programming Q&A, and technical text. It is intended for language-model pretraining
and continued pretraining. All examples use a common JSONL schema with per-record
provenance, license metadata, exact token counts, and content hashes.
The released dataset contains 2 billion tokens, 2… See the full description on the dataset page: https://huggingface.co/datasets/owenqwenllmwine/code-v1.