This dataset contains preprocessed Sundanese text data for training PixelGPT models.
Language: Sundanese (sunda)
Total samples: 294,756
Train samples: 293,933
Test samples: 823
Grapheme tokenizer: izzako/sunda-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation… See the full description on the dataset page:
https://huggingface.co/datasets/izzako/sundanese-pixelgpt.