A large-scale retrieval dataset built from LocalDoc/books_dataset — a collection of 2,804 Azerbaijani-language books with 7.8M sentences spanning politics, history, literature, science, and more. Designed for training and evaluating information retrieval, semantic search, and RAG pipelines in Azerbaijani.
The dataset consists of three configs that can be joined via passage_id and query_id:… See the full description on the dataset page:
https://huggingface.co/datasets/LocalDoc/azerbaijani_books_retriever_corpus-reranked.