This is a fine-tuned version of the multi-modal
LayoutLM model for the task of question answering on invoices and other documents. It has been fine-tuned on a proprietary dataset of
invoices as well as both
SQuAD2.0 and
DocVQA for general comprehension.
Unlike other QA models, which can only extract consecutive tokens (because they predict the start and end of a sequence), this model can predict longer-range, non-consecutive sequences with an additional
classifier head. For example, QA models often encounter this failure mode:
However this model is able to predict non-consecutive tokens and therefore the address correctly:
The best way to use this model is via
DocQuery.
This model was created by the team at
Impira.