This is a version of OpenHelix R 100k thats splitted to 14 different tasks.
The splitting has been done with all-MiniLM-L6-v2 embedding model and PCA projection.
Originally, the dataset was splitted into 16 tasks but splits having les than 3k samples (split #3 and #10) were removed as they werent necessary.
Note: Splits #3 was about generating midjourney prompts and split #10 was about really similar competitive programming problems.