This dataset contains queries and responses to evaluate financial AI chatbots for hallucinations and accuracy. The dataset was created using Lighthouz AutoBench, a no-code test case generator for LLM use cases, and then manually verified by two human annotators.
This dataset was created using Apple's 10K SEC filing from 2022. It has 100 test cases, each with a query and a response.… See the full description on the dataset page:
https://huggingface.co/datasets/lighthouzai/finqabench.