Benchmark dataset for evaluating Named Entity Recognition (NER) models on pharmaceutical packaging text.
500 synthesized pack-label texts generated from the MattBastar/Medicine_Details dataset, designed to simulate OCR output from photos of pill packaging.
Each case contains:
id: Unique case identifier
category: single_ingredient, dual_ingredient, or multi_ingredient
ocr_text: Synthesized pharmaceutical label text (clean or… See the full description on the dataset page:
https://huggingface.co/datasets/SPerva/pillchecker-ner-benchmark.