A retrieval dataset that exposes fundamental theoretical limitations of embedding-based retrieval models. Despite using simple queries like "Who likes Apples?", state-of-the-art embedding models achieve less than 20% recall@100 on LIMIT full and cannot solve LIMIT-small (46 docs).
Paper: On the Theoretical Limitations of Embedding-Based Retrieval
Code: github.com/google-deepmind/limit
Full version: LIMIT (50k documents)
Small version: LIMIT-small (46… See the full description on the dataset page:
https://huggingface.co/datasets/orionweller/LIMIT.