Dataset Card for TTC4900: A Benchmark Data for Turkish Text Categorization
Dataset Summary
The data set is taken from kemik group
The data are pre-processed for the text categorization, collocations are found, character set is corrected, and so forth.
We named TTC4900 by mimicking the name convention of TTC 3600 dataset shared by the study "A Knowledge-poor Approach to Turkish Text Categorization with a Comparative Analysis, Proceedings of CICLING 2014, Springer LNCS… See the full description on the dataset page: https://huggingface.co/datasets/savasy/ttc4900.