FinDartBench is a Korean financial question answering benchmark built from DART disclosure filings.
It is designed to evaluate real-world financial document understanding by pairing context-grounded questions with high-quality reference answers validated through a multi-stage LLM-based pipeline.
Unlike simple synthetic QA datasets, FinDartBench emphasizes grounding, answer quality, and inter-model consensus, making it suitable for reliable evaluation of financial QA… See the full description on the dataset page:
https://huggingface.co/datasets/davidkim205/FinDartBench.