PROGRESSIVE VISUAL COMPRESSION FOR EFFICIENT NAIVE-RESOLUTION ENCODING IN MLLMS
📄 Arxiv | 💻 Github
Here, we will introduce several benchmarks used in the preliminary experiments of LLaVA-UHD-v3 (ShapeGrid, Sudoku, and Sudoku in the Appendix), along with the related plotting code, preliminary experiment model inference code, and the model inference results.