Long-context tokenized corpus for benchmarking LLM prefill computation with Qwen3-8B. Contains ~10M tokens of copyright-free English text pre-tokenized with character offset mappings for fast position lookup.
data/documents.parquet
English documents with token IDs and char offsets
~100-500
data/translations.parquet
French… See the full description on the dataset page:
https://huggingface.co/datasets/di2ox3/prefill-dataset.