A large, perfectly-labeled synthetic invoice dataset for training multimodal
invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin
layouts, and all three numeral glyph systems.
Every sample is generated label-first: a structured invoice record is sampled, then
rendered to pixels via a headless browser, and bounding boxes are read back from the same
DOM — so labels and boxes are… See the full description on the dataset page:
https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.