This dataset is intended to benchmark Approximate Nearest Neighbor Search (ANNS) and Filtered Approximate Nearest Neighbor Search (FANNS) algorithms. It is based on the arXiv Dataset whose paper abstracts are embedded using the stella_en_400M_v5 embedding model. 10,000 unique arXiv search terms were generated by GPT-4 and embedded using the same Stella model to obtain the query vectors. Query attributes for… See the full description on the dataset page:
https://huggingface.co/datasets/SPCL/arxiv-for-fanns-medium.