LEAD is a synthetic training data pipeline for academic embedding models.
Quick Start
Hard Negative Sampling (Recommended for beginners)
python -c "
from beir import util
from beir.datasets.data_loader import… See the full description on the dataset page:
https://huggingface.co/datasets/LinerAI/LEAD.