Views
No views yet
infly/Infinity-Parser2-Flash, made to run on NVIDIA A100 / sm80, where the original FP8 path is unsupported. ~4.2 GB bf16 → ~2.9 GB.lm_head are kept in bf16. Calibrated on a few hundred diverse document + general-vision samples.! output. Fix: the saved config.json quantization_config.ignore uses broad regexes matching the fused names. Already applied here.| Benchmark | Published bf16 | This int4 |
|---|---|---|
| MMStar | 57.1 | 54.8 |
| OCRBench | 81.6 | 85.0 |
| DocVQA (val) | 93.2 | 93.5 |
vllm serve spectator2026/Infinity-Parser2-Flash-AWQ-W4A16 --dtype bfloat16 --trust-remote-code --reasoning-parser qwen3chat_template_kwargs={"enable_thinking": false} in requests, or answers land in the reasoning channel.