Finance-ComplexQA is a bilingual Chinese-English benchmark for complex question answering in the financial domain. It is designed to evaluate whether large language models and agent systems can answer finance questions by grounding their reasoning in reference documents rather than relying only on parametric knowledge.
The dataset covers multiple financial document domains and reasoning skills, including retrieval, multi-hop reasoning, numerical calculation… See the full description on the dataset page:
https://huggingface.co/datasets/Multilingual-Multimodal-NLP/FinanceComplexQA.