The stl_updated dataset is a large-scale collection of 3.3 million Signal Temporal Logic (STL) formulae, designed to stress-test and train models (such as Transformer encoders) on recursive understanding, semantic similarity, and syntactic complexity.
The training and test sets are augmented directly from the seed formulae introduced in Candussio (2025) and originally hosted in the base dataset saracandu/stl_formulae. These… See the full description on the dataset page: https://huggingface.co/datasets/saracandu/stl_updated.