Portuguese multi-turn conversational benchmark for measuring European and Brazilian Portuguese variety bias in LLMs.
For more details, see the P3B3 paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia… See the full description on the dataset page:
https://huggingface.co/datasets/amalia-llm/P3B3.