Mirror of the original Shitao/bge-m3-data corpus, repackaged from gzip-compressed JSONL shards into Hugging Face Dataset confi
gs stored as Parquet. Content is unchanged; this repo only standardizes the storage format and centralizes all length buckets un
der one dataset namespace.
By using this dataset, which includes up to seven hard negative texts per query (as in the original bge-m3 training data), you can easily fine-tune retrieval… See the full description on the dataset page:
https://huggingface.co/datasets/hotchpotch/bge-m3-data-finetune-unified.