SCaR-Train and SCaR-Eval are the official datasets for the paper "VIRTUE: Visual-Interactive Text-Image Universal Embedder" that are trained with MMEB-Train and SCaR-Train.
VIRTUE is a visual-interactive text-image universal embedder consisting of a VLM as well as a segmentation model to enable the visual interaction modality for human interactions.
In addition, we introduce the VIRTUE family (VIRTUE-2B-SCaR, VIRTUE-7B-SCaR), trained with MMEB-train and SCaR-Train… See the full description on the dataset page:
https://huggingface.co/datasets/Sony/SCaR-Train.