A synthetic dataset of 846 educational rewrite pairs generated for fine-tuning language models to rewrite confusing educational content in six targeted modes. Source passages were collected from Wikipedia and arXiv, and rewrites were generated using the Claude API following the Alpaca data generation methodology.
This dataset was created as part of Phase 5 of the NLP/LLM Learning Journey project and is used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/ray-2908/educational-rewriter-dataset.