This dataset is a sampled and filtered subset of the allenai/tulu-3-sft-mixture, curated and rebalanced for structured instruction fine-tuning. The goal is to support research and model development in math reasoning, coding, knowledge recall, instruction following (IF), and conversational alignment, while explicitly excluding safety, multilingual, and certain task-specific sources.
Source: Filtered from… See the full description on the dataset page:
https://huggingface.co/datasets/vanek-epfl/MNLP_M3_mcqa_dataset.