We propose VoCo-LLaMA, the first approach to compress vision tokens using LLMs. By fully utilizing the LLMs' understanding paradigm of vision tokens, our method can compress hundreds of vision tokens into a single VoCo token, while minimizing visual information loss.
VoCo-LLaMA demonstrates the ability to understand video… See the full description on the dataset page:
https://huggingface.co/datasets/lazybug/VocoUPL.