OmniRet-train is the training-data release for
OmniRet, a unified retrieval model for
text, image, video, and audio. This card documents the released snapshot for
researchers training or analyzing OmniRet.
The release contains 6,405,109 query rows and 7,119,841 candidate rows from 30
datasets. It covers 15 retrieval directions across text (T), image (I), video
(V), and audio (A). The OmniRet paper reports this corpus as… See the full description on the dataset page:
https://huggingface.co/datasets/chuonghm/OmniRet-train.