TemporalBench is a multi-domain benchmark for evaluating the temporal understanding and reasoning capabilities of large language models (LLMs) and agent-based systems over real numerical time-series.
Unlike traditional benchmarks that focus primarily on forecasting accuracy, TemporalBench is designed to diagnose how models interpret temporal structure, ground temporal patterns in context, and reason about future behavior under explicit events. To… See the full description on the dataset page:
https://huggingface.co/datasets/Melady/TemporalBench.