The first bilingual speech–document RAG dataset SV-DOC, containing Chinese and English voice queries aligned with multimodal document content, to foster future research in this direction.
SV-DOC will be released completely here. The full dataset is currently under review and will be released once approved.
To better illustrate our work, we have published 10 examples for each sub-set at github repo:… See the full description on the dataset page:
https://huggingface.co/datasets/hit12345/textlessrag.