TUS (Turkish Specialty Exam in Medicine) split of Turkish MMLU: Yapay Zeka ve Akademik Uygulamalar İçin En Kapsamlı ve Özgün Türkçe Veri Seti
This dataset is a filtered version of the original.
Embeddings with similarity score higher than 0.85 in the train split were discarded. sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 was used for this process.
Visual questions were discarded.
IMPORTANT NOTE: If you find this dataset useful, please cite the original author via Zenodo… See the full description on the dataset page:
https://huggingface.co/datasets/zypchn/TUS.