This dataset contains tokenized and packed sequences ready for LLM pretraining.
input_ids: List of token IDs (length: 2048)
attention_mask: Attention mask (1 for real tokens, 0 for padding)… See the full description on the dataset page:
https://huggingface.co/datasets/tvu-vlinhd11/pretrain-dataset-T2048-13B.