A length-bucketed release of togethercomputer/Long-Data-Collections, used in Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference (ICML 2026 Spotlight).
Each example is assigned to exactly one length bucket based on token count with meta-llama/Llama-3.1-8B-Instruct (add_special_tokens=True):
32k_64k
[32K, 64K)… See the full description on the dataset page:
https://huggingface.co/datasets/sxiong/DHSA_Long-Data-Collections.