Reads a receipt image and returns it as structured JSON. Full-precision master: the merged bf16 model every other artifact in this project is derived from.
Turning a photo or scan of a retail receipt into structured JSON, for bookkeeping and expense
pipelines, or as a starting point for research on document understanding.
Out of scope
Invoices, handwritten notes, and non-receipt images. The model was never trained on them. It
also should not be trusted to make financial decisions on its own; treat the output as an
extraction to be reviewed.
How to use
The model expects one image and this exact prompt, which is the prompt it was trained with:
Extract the receipt from the image into a structured JSON. Your output should contain ONLY correct JSON!
python
1import torch
2from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
34REPO ="<your-hf-username>/qwen3vl-4b-receipt-extraction"5model = Qwen3VLForConditionalGeneration.from_pretrained(REPO, dtype=torch.bfloat16, device_map="auto")6processor = AutoProcessor.from_pretrained(REPO, max_pixels=1600*28*28)78messages =[{"role":"user","content":[9{"type":"image","image":"receipt.jpg"},10{"type":"text","text":"Extract the receipt from the image into a structured JSON. Your output should contain ONLY correct JSON!"},11]}]12inputs = processor.apply_chat_template(13 messages, add_generation_prompt=True, tokenize=True,14 return_dict=True, return_tensors="pt",15).to(model.device)1617out = model.generate(**inputs, max_new_tokens=512, do_sample=False)18print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Output
A single JSON object with three top-level keys: menu (the line items), sub_total and
total. Every value is a string, copied from the receipt.
json
1{"menu":[{"nm":"Es Teh Manis","cnt":"2","price":"8,000"}],2"sub_total":{"subtotal_price":"8,000"},3"total":{"total_price":"8,000","cashprice":"10,000","changeprice":"2,000"}}
Evaluation
CORD-v2 test split, 100 receipts, greedy decoding, images capped at 1600 patches.
field_f1
nTED
json_validity
exact_match
0.8862
0.9105
0.9900
0.4500
Measured in results/finetune-result.json (eval_script.py, plain decode).
field_f1 flattens the predicted and gold JSON into (field-path, value) leaf pairs and
scores precision/recall over them. It asks whether the right values landed at the right
paths.
nTED (normalized tree-edit distance, following the CORD paper) compares the two JSON
trees directly. It catches structural mistakes, like wrong nesting or a missing key, that
leaf-flattening misses.
Neither subsumes the other, which is why both are here.
Training
Fine-tuned from Qwen/Qwen3-VL-4B-Instruct with LoRA on naver-clova-ix/cord-v2
(800 train / 100 validation / 100 test receipts). Labels were losslessly normalized before
training.
LoRA rank / alpha / dropout
16 / 32 / 0.05
Adapted modules
language tower only (q,k,v,o,gate,up,down_proj); vision tower frozen
Precision
bf16
Learning rate
1e-4
Batch size
2 per device x 4 gradient accumulation
Epochs
20
Image budget
1600 patches of 28x28
Training ran well past convergence on purpose, so there was a wide range of adapters to pick
from. Validation loss and the task metrics peak at different times here: loss starts rising
around epoch 3, while the best field_f1 and nTED land at epoch 9 or later. The adapter
published here is the one that scored best on the validation split by field_f1, and it was
then scored once on the held-out test split. Early-stopping on loss would have picked a
noticeably worse one.
Limitations
The test split is 100 receipts. Gaps of a point or two in field_f1 are inside the noise band
and should be read as ties.
CORD-v2 is photographed Indonesian retail and restaurant receipts. Expect worse results on
other layouts, languages, or document types (invoices, handwriting).
Field names mirror the raw CORD keys (nm, cnt, unitprice), not friendly names.
Values are transcribed text, not validated arithmetic. Nothing checks that the line items
sum to the total. Do not use the output for financial decisions without review.
Schema-constrained decoding was tried and scored far worse than plain decoding (field_f1
falls roughly 35 points), so plain greedy decoding is what these numbers use and what is
recommended.
Any timing or throughput figure depends heavily on hardware. Measure on the machine you
intend to deploy on.
License
apache-2.0, inherited from the base model. The CORD-v2 dataset carries its own terms; see
its dataset card.