ROSIE-MIND-Topics is a bilingual (English–Spanish) dataset of 875,230 passages containing topic modeling information derived from training a PLTM model on this data with 30 topics. It serves as the input dataset for the MIND pipeline, which detects multilingual and cultural discrepancies in question–answer pairs.
Each record includes the passage and corresponding full document, preprocessing outputs (lemmas, translations), and topic model features (topic… See the full description on the dataset page:
https://huggingface.co/datasets/lcalvobartolome/rosie_mind_topics.