This dataset derives deterministic, mutually disjoint splits from
kushasareen/sdpo-datasets/polaris_5k
at source commit 51fd81230be64a06d77e7edc33d3c0d60b3186b5.
train (train_split.parquet)
51,911
RL actor training and GEPA trajectory training
test (test.parquet)
500
Final model evaluation… See the full description on the dataset page:
https://huggingface.co/datasets/Cameron-Chen/polaris-53k.