STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure.
Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity:
State Estimation:… See the full description on the dataset page:
https://huggingface.co/datasets/prquan/STARK_10k.