This dataset contains preprocessed and chunked Wikipedia HTML dumps from 25 languages.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources with Retrieval and Semantic Parsing
WikiChat
Stopping the… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia.