An open-source, fully-documented dataset (and eventually training pipeline)
for fine-tuning small and large agentic coding models on real, extracted
engineering work — not synthetic toy problems.
This repo is public end-to-end: the dataset, the conversion tooling, and
(soon) the actual fine-tuning code will all live here or in the linked
GitHub repo, so anyone can inspect exactly how the data was built and how
the models were… See the full description on the dataset page:
https://huggingface.co/datasets/sinamsv00/Luna.