Arabic query-product training data for fine-tuning retrieval and embedding models on e-commerce
catalog search in Modern Standard Arabic and Libyan dialect.
This public dataset exposes only query text and product-title text.
Evaluation benchmark: this is the training counterpart to
prestoai/arabic-ecom-search-bench.
Train here, evaluate there.
Each subset ships an explicit train/test split.
Subset
Train
Test… See the full description on the dataset page:
https://huggingface.co/datasets/prestoai/arabic-ecom-data.