This repository contains k-means cluster subsets for training. Each subset is a config (name=cluster_X), backed by Parquet.
ds = load_dataset("{username}/fykcluster-k4", data_files={"train":… See the full description on the dataset page:
https://huggingface.co/datasets/Adanato/fykcluster-k4.