The SynthSTEL dataset is a synthetically generated (with GPT-4) extension of the STEL task to 40 style features. This data was also used to train the StyleDistance embedding model.
This synthetic dataset was produced with DataDreamer 🤖💤.
This research is supported in part by the Office of the Director of National Intelligence (ODNI), Intelligence Advanced Research Projects Activity (IARPA), via the HIATUS Program… See the full description on the dataset page:
https://huggingface.co/datasets/StyleDistance/synthstel.