This dataset is part of the CS-552 Modern NLP (Spring 2025) course project at EPFL. It is the tokenized and instruction-formatted version of the MNLP_M2_quantized_dataset, specifically prepared for training quantized large language models (LLMs) using techniques such as QLoRA and LoRA fine-tuning.
The dataset is optimized for causal language modeling (CLM) tasks and intended to:
Support training of 4-bit quantized models… See the full description on the dataset page:
https://huggingface.co/datasets/abdou-u/MNLP_M2_quantized_dataset_formatted.