Dataset type:
TaiVision-pretrain-1M-v1.0 is a traditional Chinese Image Captioning dataset. This dataset is a concatenation of the two datasets: liuhaotian/LLaVA-Pretrain and liuhaotian/LLaVA-CC3M-Pretrain-595K translated into traditional Chinese by yentinglin/Llama-3-Taiwan-8B-Instruct.
License:
Must comply with license of CC-3M, BLIP (if you use their synthetic caption).