Public dataset for BizBench.
Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge.
Together, these requirements make this domain difficult for large language models (LLMs).
We introduce BizBench, a benchmark for evaluating models' ability to reason about realistic financial problems.
BizBench comprises eight quantitative reasoning tasks… See the full description on the dataset page:
https://huggingface.co/datasets/kensho/bizbench.