Turning a photo or scan of a retail receipt into structured JSON, for bookkeeping and expense
pipelines, or as a starting point for research on document understanding.
Out of scope
Invoices, handwritten notes, and non-receipt images. The model was never trained on them. It
also should not be trusted to make financial decisions on its own; treat the output as an
extraction to be reviewed.
How to use
The model expects one image and this exact prompt, which is the prompt it was trained with:
Extract the receipt from the image into a structured JSON. Your output should contain ONLY correct JSON!
Needs GPTQModel (plain transformers would need optimum).
python
1import torch
2from gptqmodel import GPTQModel
3from transformers import AutoProcessor
45REPO ="<your-hf-username>/qwen3vl-4b-receipt-extraction-gptq"6model = GPTQModel.load(REPO, device_map="auto").model # .model is the underlying HF model7processor = AutoProcessor.from_pretrained(REPO, max_pixels=1600*28*28)89messages =[{"role":"user","content":[10{"type":"image","image":"receipt.jpg"},11{"type":"text","text":"Extract the receipt from the image into a structured JSON. Your output should contain ONLY correct JSON!"},12]}]13inputs = processor.apply_chat_template(14 messages, add_generation_prompt=True, tokenize=True,15 return_dict=True, return_tensors="pt",16).to(model.device)1718out = model.generate(**inputs, max_new_tokens=512, do_sample=False)19print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
It also serves under vLLM, which reads the same GPTQ weight format.
Output
A single JSON object with three top-level keys: menu (the line items), sub_total and
total. Every value is a string, copied from the receipt.
json
1{"menu":[{"nm":"Es Teh Manis","cnt":"2","price":"8,000"}],2"sub_total":{"subtotal_price":"8,000"},3"total":{"total_price":"8,000","cashprice":"10,000","changeprice":"2,000"}}
Evaluation
CORD-v2 test split, 100 receipts, greedy decoding, images capped at 1600 patches.
field_f1
nTED
json_validity
exact_match
0.8840
0.9143
0.9800
0.4100
Measured in section 7.2 of finetune_qwen3vl.ipynb.
field_f1 flattens the predicted and gold JSON into (field-path, value) leaf pairs and
scores precision/recall over them. It asks whether the right values landed at the right
paths.
nTED (normalized tree-edit distance, following the CORD paper) compares the two JSON
trees directly. It catches structural mistakes, like wrong nesting or a missing key, that
leaf-flattening misses.
Neither subsumes the other, which is why both are here.
Training
Fine-tuned from Qwen/Qwen3-VL-4B-Instruct with LoRA on naver-clova-ix/cord-v2
(800 train / 100 validation / 100 test receipts). Labels were losslessly normalized before
training.
LoRA rank / alpha / dropout
16 / 32 / 0.05
Adapted modules
language tower only (q,k,v,o,gate,up,down_proj); vision tower frozen
Precision
bf16
Learning rate
1e-4
Batch size
2 per device x 4 gradient accumulation
Epochs
20
Image budget
1600 patches of 28x28
Training ran well past convergence on purpose, so there was a wide range of adapters to pick
from. Validation loss and the task metrics peak at different times here: loss starts rising
around epoch 3, while the best field_f1 and nTED land at epoch 9 or later. The adapter
published here is the one that scored best on the validation split by field_f1, and it was
then scored once on the held-out test split. Early-stopping on loss would have picked a
noticeably worse one.
Quantization
Quantized with GPTQModel (4-bit, group_size=128, symmetric, desc_act=false), which works
out to roughly 4.29 bits per weight. Only the language-model decoder linears are quantized;
the vision tower, the embeddings and lm_head stay in bf16.
Calibration used 256 multimodal chat samples from the CORD-v2 train split, each one a receipt
image plus the fixed prompt plus the gold JSON answer. That makes this a data-aware
quantization, and it shows: against the bf16 model it loses 0.0022 field_f1 and actually gains
0.0038 nTED, which is inside the noise of a 100-receipt split. A data-free bitsandbytes NF4
quantization of the same model came out measurably worse on both metrics.
Inference runs through the Marlin 4-bit kernel.
Limitations
The test split is 100 receipts. Gaps of a point or two in field_f1 are inside the noise band
and should be read as ties.
CORD-v2 is photographed Indonesian retail and restaurant receipts. Expect worse results on
other layouts, languages, or document types (invoices, handwriting).
Field names mirror the raw CORD keys (nm, cnt, unitprice), not friendly names.
Values are transcribed text, not validated arithmetic. Nothing checks that the line items
sum to the total. Do not use the output for financial decisions without review.
Schema-constrained decoding was tried and scored far worse than plain decoding (field_f1
falls roughly 35 points), so plain greedy decoding is what these numbers use and what is
recommended.
Any timing or throughput figure depends heavily on hardware. Measure on the machine you
intend to deploy on.
License
apache-2.0, inherited from the base model. The CORD-v2 dataset carries its own terms; see
its dataset card.