This dataset is part of the CS-552 Modern NLP (Spring 2025) course project at EPFL. It contains a merged and cleaned collection of multiple-choice question-answer (MCQA) datasets curated for training generative reasoning models.
The dataset was constructed to support training and evaluation of large language models (LLMs) on complex multi-step reasoning tasks, including those from:
Algebra and Arithmetic Word Problems
Natural… See the full description on the dataset page:
https://huggingface.co/datasets/abdou-u/MNLP_M2_quantized_dataset.