Latest exported JSON files for the multimodal information-retrieval project.
multimodal_documents.g4.captioned.s2tw.json
769,245
Latest document corpus with image captions and Simplified-to-Traditional Chinese conversion
e4e3f7609fde3297d89e436a11751dc5a394109f7aba2ae487fe9323d198802d
multimodal_pretrain_pairs.json
351,979
Latest large pretraining/query pairs with text… See the full description on the dataset page:
https://huggingface.co/datasets/Sigoso12/multimodal-ir-zh-tw.