This dataset repository contains Turkish supervised fine-tuning datasets prepared for the NEDO Turkish SLM project.
The datasets were built to fine-tune NEDOQwen-style Turkish decoder-only language models after base pretraining on the NEDO Turkish 65K tokenized web corpus.
Main related pretraining dataset:
Ethosoft/nedo-turkish-65k-tokenized-60b
File
Examples
Recommended?
Description… See the full description on the dataset page:
https://huggingface.co/datasets/Ethosoft/nedo-turkish-sft-mixtures.