The dataset is stored at the OSF here
MLRegTest is a benchmark for sequence classification, containing training, development, and test sets from 1,800 regular languages.
Regular languages are formal languages, which are sets of sequences definable with certain kinds of formal grammars, including
regular expressions, finite-state acceptors, and monadic second-order logic with either the successor or precedence relation in the
model signature for words. This benchmark was designed to help… See the full description on the dataset page:
https://huggingface.co/datasets/samvdp/MLRegTest.