The original HANNA dataset (Chhun et al., 2022) contains 1,056 stories, each annotated by human raters using a 5-point Likert scale across six criteria: Relevance, Coherence, Empathy, Surprise, Engagement, and Complexity. These stories are based on 96 story prompts from the WritingPrompts dataset (Fan et al., 2018), with each prompt generating 11 stories, including one human-written and 10 generated by different automatic text generation… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/hanna-annotated-latest.