GRAS (Grading at Scale) is a semi-synthetic dataset for automatic grading of short answers (ASAG) using large language models (LLMs).
This dataset contains student answers to questions across four domains (Neuroscience, Psychology and AI), with labels indicating whether each answer is correct, partially correct, or incorrect.
The student answers are synthetically generated with GPT-4o.
Splits: train… See the full description on the dataset page:
https://huggingface.co/datasets/saurluca/GRAS.