The MMLU dataset's validation split was filtered to include only subjects categorized as STEM and health, following the official subject taxonomy provided in the MMLU GitHub repository and as described in Hendrycks et al. (2021).
All selected STEM and health subject test sets were merged, randomly shuffled with a fixed seed, and assigned unique IDs to create a unified calibration set, as outlined in the original MMLU paper: "Measuring Massive… See the full description on the dataset page:
https://huggingface.co/datasets/casimiir/MNLP_M3_quantized_dataset.