collection of several text2image prompt datasets
data was cleaned/normalized with the goal of removing "model specific APIs" like the "--ar" for Midjourney and so on
data de-duplicated on a basic level: exactly duplicate prompts were dropped (after cleaning and normalization)