Training dataset for the distill-expert model — a 0.6B LoRA fine-tuned from Qwen3-0.6B that compresses shell/command output for AI consumption.
train.jsonl.gz — 454,710 training examples (85%)
valid.jsonl.gz — 53,458 validation examples (10%)
test.jsonl.gz — 26,832 test examples (5%)
runpod_train.py — Unsloth LoRA training script (RunPod-ready)
Total: 535,000 examples across 8 operation modes.
Mode
Examples… See the full description on the dataset page:
https://huggingface.co/datasets/samuelfaj/distill-expert-535k.