Supervised fine-tuning mixture for improving visual table question answering in
small vision-language models — built for the MiniVLMDocEval project to lift
Qwen3.5-0.8B on TableVQABench.
We measured Qwen3.5-0.8B per TableVQABench sub-domain and found the weakness is
Wikipedia-style visual-table lookup, not financial tables:
vwtq (Wikipedia lookup)
27.8
weakest, and… See the full description on the dataset page:
https://huggingface.co/datasets/savoji/minivlm-tablevqa-sft.