AraReceipt is a manually annotated dataset of 100 Arabic retail receipt images labeled with 25 key-information classes plus an Ignore category, in WildReceipt style. It contains 4,609 annotated regions (46.1 ± 21.6 per image), each with a box, a transcription, and a semantic class.
The dataset was built with GAIDA, a human-in-the-loop annotation system that combines OCR-based region extraction, LLM-based semantic pre-annotation, and human validation in Label Studio.… See the full description on the dataset page:
https://huggingface.co/datasets/IslamMesabah/AraReceipt.