This dataset is a lightweight, filtered version of the original Jackrong/DeepSeek-V4-Distill-8000x dataset.
It has been specifically filtered to strictly include samples with a maximum token count of 2048.
Why use this dataset?
The original distillation dataset contains incredibly rich reasoning traces, but many of them are exceptionally long. If you are fine-tuning on consumer hardware (like RTX… See the full description on the dataset page: https://huggingface.co/datasets/khang272/DeepSeek-V4-Distill-8000x-2048.