Dataset Card for llm-slice/storytelling_anthology
Dataset Summary
This dataset is an anthology of short story completions generated by a series of language model checkpoints during interactive reinforcement learning with a storytelling objective. Each branch (e.g., chck_20M, chck_90M, chck_900M) corresponds to models pretrained on increasing numbers of words, then further trained using Proximal Policy Optimization (PPO) against a teacher model (Llama 3.1-8B-Instruct).
The… See the full description on the dataset page: https://huggingface.co/datasets/llm-slice/storytelling_anthology.