This dataset contains preprocessed data for diffusion language model training.
Padded example num: 97
Truncated example num: 903
Have edit tokens example num: 1000
Edit token ratio: 0.1576+/-0.0246
Padding length: 1330.7732+/-838.2759
Maximum sequence length: 4096
System prompt: If you are not confident about a token prediction, output the exact placeholder <|reserved_token_0|>… See the full description on the dataset page:
https://huggingface.co/datasets/s-ryoma/s1k-random-15percent-edit-4096-padding.