LongDS-Bench is a benchmark for evaluating long-horizon, multi-turn agentic data analysis. Real-world analysis is rarely a sequence of independent questions: filters, metric definitions, assumptions, intermediate tables, and branch-specific results evolve over many turns. LongDS tests whether agents can maintain and apply these evolving analytical states correctly.
LongDS contains 68 tasks constructed… See the full description on the dataset page:
https://huggingface.co/datasets/zjunlp/LongDS.