VLE (Visual-Language Encoder) is an image-text multimodal understanding model built on the pre-trained text and image encoders.
It can be used for multimodal discriminative tasks such as visual question answering and image-text retrieval.
Especially on the visual commonsense reasoning (VCR) task, which requires high-level language understanding and reasoning skills, VLE achieves significant improvements.