A finance-domain dataset for evaluating retrieval, reranking, and RAG systems under realistic and challenging conditions.
⚠️ This dataset is intentionally low-overlap.High performance from keyword-based methods (e.g., BM25) likely indicates shortcut exploitation rather than real semantic understanding.
This dataset's queries were generated using gpt-oss-120b, served via regolo.ai.… See the full description on the dataset page:
https://huggingface.co/datasets/ReDiX/ReDiX-Benchmark-Finance-ita.