Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks.
The main purpose of this dataset is to provide a practical, scalable, and⦠See the full description on the dataset page:
https://huggingface.co/datasets/Ethosoft/Turkish_corpus.