RoleTriviaQA is the dataset used in the Role-Playing Knowledge Speech QA task in the paper What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study.
RoleTriviaQA is an open-source General Role-Playing Knowledge Speech QA dataset. The dataset contains 138,384 train, 300 validation and 2,426 test samples based on the original TriviaQA dataset with 15 characterized voices from roles in Genshin… See the full description on the dataset page:
https://huggingface.co/datasets/cnxup/RoleTriviaQA.