Personal Codex Model Training Corpus is a provenance-aware, repository-level dataset for causal
language modeling, code completion, continued pretraining, and coding assistant adaptation. It is
built from source files present in local Git repository checkouts at a defined collection point.
The dataset prioritizes broad, authentic software-engineering coverage while retaining enough
metadata to audit every emitted… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/personal-codex-model.