LongFinanceQA dataset is designed to generate practical long-context QA pairs with reasoning steps to effectively analyze long content. It consists of 46,457 long-context QA pairs.
Paper and more resources: [arXiv] [Project Website]
This dataset is used for academic research purposes only.
Below is a sample from the dataset:
{
"id": "id_000000",
"doc": "docs/doc_0800.txt",
"question":… See the full description on the dataset page:
https://huggingface.co/datasets/jylins/LongFinanceQA.