This dataset contains preprocessed data for diffusion language model training.
Padded example num: 115
Truncated example num: 635
Have edit tokens example num: 750
Edit token ratio: 0.1671+/-0.0211
Maximum sequence length: 256
System prompt: If you are not confident about a token prediction, output the exact placeholder <|reserved_token_0|> instead of that token. Generate… See the full description on the dataset page:
https://huggingface.co/datasets/s-ryoma/lima-filtered-random-15percent-edit-256-padding.