Final cleaned and oversampled dataset for SFT training on Qwen3.5-4B
for the EBOS medical document extraction task.
Broken instruction schema fixed — removed extra closing brace in patient block
MedicareNumber removed from instruction — zero examples in dataset, removed from schema
31 null output rows dropped — rows with no ground truth removed
Medicine names normalized to title case — BUPRENORPHINE → Buprenorphine… See the full description on the dataset page:
https://huggingface.co/datasets/abideen/medicine-2.