Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1
(High-Quality-DQA subset) and tokenized with the Qwen/Qwen3.6-27B tokenizer (vocab 248,320).
This is a re-tokenization of the same corpus previously released with the Qwen/Qwen3-8B
tokenizer. Qwen3.6 uses a different, larger vocabulary, so the old token ids are not valid
for Qwen3.6 models — the QA pairs were re-extracted from the raw… See the full description on the dataset page:
https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3.6-tokens.