This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately.
queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page:
https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.