Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models)
Total chunks: ~2,520,000… See the full description on the dataset page:
https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.