COCO-CN is a bilingual image description dataset, developed based on (a subset of) MS-COCO with manually written Chinese sentences and tags.
The dataset can be used for multiple tasks including image tagging, captioning and retrieval, all in a cross-lingual setting.
Chinese and English