This dataset is a Vietnamese triplet dataset for dense retrieval training.
It combines Vietnamese text from unicamp-dl/mmarco with hard-negative passage ids from sentence-transformers/msmarco-hard-negatives.
bm25_rank1_10_1m
1,000,000
bm25
2 per positive
1-10… See the full description on the dataset page:
https://huggingface.co/datasets/QuangDuy/mmarco-vi-hard-negatives.