This Language Identification Dataset provides a multi-domain corpus in European and Brazilian Portuguese.
The repository is an anonymyzed version to support a submsission to the EACL 2024 conference.
Further information about the dataset can be soon found in the paper: Enhancing Portuguese Variants Identification with Domain-Agnostic Ensemble Approaches