This repository contains k-means cluster subsets for training. Each subset is a config (name=cluster_X), backed by Parquet.
ds = load_dataset("{username}/10model_rank_k3"… See the full description on the dataset page:
https://huggingface.co/datasets/Adanato/10model_rank_k3.