PROGRESSIVE VISUAL COMPRESSION FOR EFFICIENT NAIVE-RESOLUTION ENCODING IN MLLMS
📄 OpenReview | 💻 Github
Here, we will introduce several benchmarks used in the preliminary experiments of LLaVA-UHD-v3 (ShapeGrid, Sudoku, and Sudoku in the Appendix), along with the related plotting code, preliminary experiment model inference code, and the model inference results.
The pilot experiment is designed to… See the full description on the dataset page:
https://huggingface.co/datasets/testeam/Pilot_experiment.