This dataset was generated using YourBench (v0.9.0), an open-source framework for generating domain-specific benchmarks from document collections.
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based… See the full description on the dataset page:
https://huggingface.co/datasets/jenny0830/celarai_early_literacy_public_Llama-31-8B-Instruct.