This dataset is a pre-tokenized, shuffled, and interleaved version of all subsets from turkish-nlp-suite/ForumSohbetleri.
Tokenizer: Ba2han/qwen-test-3
Filtering: Min tokens = 50, Max tokens = 2550
Format: uint32 input_ids only.
Subsets Included: donanimarsivi, donanimhaber, forumum, iyinet, kadinlarklubu, memurlar, tahribat, technopatsosyal, turkiyeforum, wardom, wmaraci
Shuffling: Stream interleaved with a… See the full description on the dataset page:
https://huggingface.co/datasets/Ba2han/forumsoh_tokenized.