Invoice images with structured ground-truth output, formatted for document-understanding training. Directly the shape of our core product problem.
We use it for: fine-tuning and evaluating invoice extraction - validating that schema-constrained decoding holds up on real layouts.
This is an unmodified fork of katanaml-org/invoices-donut-data-v1, created by the Qwen team.
All… See the full description on the dataset page:
https://huggingface.co/datasets/NeuralMetrics/invoices-donut-data-v1.