This synthetic dataset is designed for training and evaluating Information Retrieval (IR) models, especially for document relevance classification tasks. It contains query-document pairs with binary relevance labels.
Queries and documents are generated for three distinct topics:
sports: football, basketball, tennis, match, score
tech: AI, machine learning, cloud, software, algorithm
health: diet, exercise, nutrition, wellness, fitness… See the full description on the dataset page: https://huggingface.co/datasets/Karimansour/Synthetic_IR.