NVFP4 (4-bit float) quantization of baidu/Unlimited-OCR,
a 3B vision-language OCR model that pushes DeepSeek-OCR one step further (one-shot,
long-horizon document parsing). This repo quantizes the DeepSeek-V2 MoE text decoder to
NVFP4 while keeping the vision tower in BF16, so it stays a drop-in transformers model.
⚠️ Runtime requirements. This is custom remote code, so load with
trust_remote_code=True, transformers 4.57.x, and compressed-tensors installed.
NVFP4 runs natively on Blackwell GPUs (Jetson Thor, RTX 50-series, B200); on other GPUs
compressed-tensors transparently dequantizes the weights at load.
This quant
Scheme
NVFP4A16 · 4-bit float · group 16 · nvfp4-pack-quantized
Unlimited-OCR uses the DeepSeek-OCR prompt vocabulary. The prompt must contain <image>;
prefix it with <|grounding|> whenever you also want bounding boxes for what was read.
Task
Prompt
Document → Markdown (layout-aware, with boxes)
`
\n<
Plain text OCR (just the text, no layout)
<image>\nFree OCR.
OCR with bounding boxes
`
\n<
Native Unlimited-OCR parse
<image>document parsing.
Parse a figure / chart / diagram
<image>\nParse the figure.
Describe the image (general VQA)
<image>\nDescribe this image in detail.
Find specific text (referring grounding)
`
\n<
Multi-page / PDF
<image>Multi page parsing. via model.infer_multi(...)
Resolution modes
base — base_size=1024, image_size=1024, crop_mode=False. Good default for normal pages.
gundam — base_size=1024, image_size=640, crop_mode=True. Tiles the page; use for dense
or large/high-resolution documents.
Understanding the output (grounding tokens)
With <|grounding|>, the model interleaves the recognized text with detection boxes:
Each [x1, y1, x2, y2] is the bounding box (top-left → bottom-right) of that span, in the
coordinate space of the model's input image. Drop the <|det|>...<|/det|> tags if you only want
text, or parse them to overlay boxes / rebuild layout. Without <|grounding|> you get plain text
(or Markdown) with no box tags.
Serving
The original model ships an SGLang wheel and a vLLM path (see the
base model card). For quantized serving, a runtime
with compressed-tensors support can load the NVFP4 weights directly; otherwise use the
transformers snippet above.
Task: multilingual OCR / document parsing — single image, multi-page, and PDF
(one-shot long-horizon parsing).
License: MIT (inherited from the base model).
How this was made
NVFP4 was applied with llm-compressor's model_free_ptq — a data-free path that streams
the safetensors and quantizes weights tensor-by-tensor (no calibration, no model forward), so the
custom VLM code is irrelevant. The vision tower, projector, embeddings, lm_head, MoE router and
norms were excluded via ignore patterns and remain BF16.
Verified
Loaded in transformers and run on a test document — OCR output is identical to BF16, e.g.:
Very-low-bit weight quant trades a little accuracy for size; for the highest fidelity use the
original BF16 model. For OCR, NVFP4 here is effectively lossless on tested documents.
The vision encoder stays BF16 regardless (small, and accuracy-sensitive).
English-/multilingual-text centric; verify critical fields on hard scans.