~7,000 chat examples for sequence-level (data) distillation. ProofKit's fine-tuned
gpt-oss-20b teacher (visproj/proofkit-gpt-oss-20b-lora)
regenerates the assistant turn over the exact prompts from
visproj/proofkit-sft; the
system + user turns are kept verbatim, so the set stays license-safe (no scraping, no
PII).
A Qwen 0.5B student is then SFT'd on this to produce
visproj/proofkit-distilled-qwen0.5b
(and its GGUF), which the ProofKit Space serves… See the full description on the dataset page:
https://huggingface.co/datasets/build-small-hackathon/proofkit-distill-qwen0.5b.