TexAes is the first aesthetic dataset in the LLM domain, containing a total of 50,390 prompts. It is curated by an aesthetic data generation pipeline leveraging GPT-4o for aesthetic polishing, as described in our paper "Textual Aesthetics in Large Language Models."
To address the challenge of generating high-quality aesthetic preference data, we developed a scalable aesthetic data generation pipeline. This pipeline utilizes GPT-4o to enhance the aesthetic… See the full description on the dataset page:
https://huggingface.co/datasets/lingjie23/TexAes.