This dataset was used in the paper GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning. Refer to
https://arxiv.org/abs/2402.16829 for details.
The code for generating the data is available at
https://github.com/avsolatorio/GISTEmbed/blob/main/scripts/create_classification_dataset.py.
@article{solatorio2024gistembed,
title={GISTEmbed: Guided In-sample Selection of Training… See the full description on the dataset page:
https://huggingface.co/datasets/avsolatorio/mteb-amazon_massive_scenario-avs_triplets.