QAmeleon introduces synthetic multilingual QA data contaning in 8 langauges using PaLM-540B, a large language model. This dataset was generated by prompt tuning PaLM with only five examples per language. We use the synthetic data to finetune downstream QA models leading to improved accuracy in comparison to English-only and translation-based baselines.
Data available at
https://storage.googleapis.com/qameleon/qamelon_pt_accepted.csv
More details can be… See the full description on the dataset page:
https://huggingface.co/datasets/imvladikon/QAmeleon.