Does a fine-tuned 1.5B extraction model hallucinate more than the Qwen-7B it
replaced? This dataset is the measurement: three systems, 25 held-out
transcripts, two groundedness metrics, with the raw generations included so
anyone can re-score with their own method.
prompted Qwen-7B — the original feature extractor
prompted Qwen-1.5B base — same size as the LoRA, no fine-tune (control)
fine-tuned 1.5B LoRA — the shipped… See the full description on the dataset page:
https://huggingface.co/datasets/legendarydragontamer/turboskillslug-groundedness-eval.