Italian Synthetic Retrieval Dataset with Dense Hard Negatives
Dataset Summary
This dataset is a high-quality, synthetic Information Retrieval (IR) dataset for the Italian language. It is designed to train and fine-tune state-of-the-art embedding models and bi-encoders (e.g., using MultipleNegativesRankingLoss).
The dataset consists of exactly 50,000 synthetic search queries generated from approximately 25,000 unique passages extracted from the Italian Wikipedia.… See the full description on the dataset page: https://huggingface.co/datasets/nickprock/it-wiki-retrieval-synthetic-hn.