This dataset is associated with the paper KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search.
It contains 3,000 cleaned full-text Khmer web documents, 300 manually reviewed Khmer search queries, and silver relevance labels. The dataset is designed for sparse, dense, hybrid, and LLM-assisted retrieval experiments.
The repository provides the following files:… See the full description on the dataset page:
https://huggingface.co/datasets/Backkh/KSE-Web3K.