Figure 1: Domain/task distribution and task illustrations in TSAQA.
TSAQA is a large-scale, unified benchmark for evaluating the temporal analytical capabilities of language models on time series data. It addresses the fragmented landscape of existing time series QA benchmarks by consolidating 6 diverse analytical tasks under a single standardized evaluation framework.
The… See the full description on the dataset page:
https://huggingface.co/datasets/TSAQA/TSAQA-Benchmark.