This repository contains the dataset used for pretraining the Turkish-LLaVA-v0.1 model. The dataset is a Turkish translation of the English dataset used in previous studies liuhaotian. The translation was performed using DeepL. The details of this dataset and its comparison with other datasets have been published in our paper (Soon..).