Views
No views yet


| Benchmarks | EMOVA-3B | EMOVA-7B | EMOVA-72B | GPT-4o | VITA 8x7B | VITA 1.5 | Baichuan-Omni |
|---|---|---|---|---|---|---|---|
| MME | 2175 | 2317 | 2402 | 2310 | 2097 | 2311 | 2187 |
| MMBench | 79.2 | 83.0 | 86.4 | 83.4 | 71.8 | 76.6 | 76.2 |
| SEED-Image | 74.9 | 75.5 | 76.6 | 77.1 | 72.6 | 74.2 | 74.1 |
| MM-Vet | 57.3 | 59.4 | 64.8 | - | 41.6 | 51.1 | 65.4 |
| RealWorldQA | 62.6 | 67.5 | 71.0 | 75.4 | 59.0 | 66.8 | 62.6 |
| TextVQA | 77.2 | 78.0 | 81.4 | - | 71.8 | 74.9 | 74.3 |
| ChartQA | 81.5 | 84.9 | 88.7 | 85.7 | 76.6 | 79.6 | 79.6 |
| DocVQA | 93.5 | 94.2 | 95.9 | 92.8 | - | - | - |
| InfoVQA | 71.2 | 75.1 | 83.2 | - | - | - | - |
| OCRBench | 803 | 814 | 843 | 736 | 678 | 752 | 700 |
| ScienceQA-Img | 92.7 | 96.4 | 98.2 | - | - | - | - |
| AI2D | 78.6 | 81.7 | 85.8 | 84.6 | 73.1 | 79.3 | - |
| MathVista | 62.6 | 65.5 | 69.9 | 63.8 | 44.9 | 66.2 | 51.9 |
| Mathverse | 31.4 | 40.9 | 50.0 | - | - | - | - |
| Librispeech (WER↓) | 5.4 | 4.1 | 2.9 | - | 3.4 | 8.1 | - |
1@article{chen2024emova,
2 title={Emova: Empowering language models to see, hear and speak with vivid emotions},
3 author={Chen, Kai and Gou, Yunhao and Huang, Runhui and Liu, Zhili and Tan, Daxin and Xu, Jing and Wang, Chunwei and Zhu, Yi and Zeng, Yihan and Yang, Kuo and others},
4 journal={arXiv preprint arXiv:2409.18042},
5 year={2024}
6}