This dataset was developed to support research in ad matching, semantic retrieval, and intent alignment across offering and wanted ads in Sri Lankan classified marketplaces. It consists of 54,489 ad pairs.
It includes both human-verified real samples and LLM-generated synthetic samples.
Designed for training and evaluating ML models requiring generalization across low-resource subcategories
Especially valuable where wanted ads are… See the full description on the dataset page:
https://huggingface.co/datasets/Damika-7/sri_Lankan_classified_ads_dataset_for_ad_matching_and_retrieval.