Portuguese preprocessed split from MQA dataset containing only the question_title and answer_text columns of records in the ".pt" domain.
The dataset was derived by filtering the following dataset: ju-resplande/qa-pt
The rationale is to have a dataset that is closer aligned with the Portuguese (Portugal) language.
Semantic deduplication splits included for thresholds of 0.7, 0.8 and 0.9 with model… See the full description on the dataset page:
https://huggingface.co/datasets/marquesafonso/qa-ptpt.