I sampled ~50k instances from
https://beaker.org/ds/01J7WZNKXSKRJJYQ1P0H9JFW04/details.
I tried to sample equally across datasets.
Useful for some preference experiments.
Looking for shards? (213 shards with 250 rows each) 💎
Beaker: ljm/tulu3.4-sft-replica-50k-shards
AWS: s3://ai2-ljm-dev/tulu3.4-sft-replica-50k/*.jsonl
Looking for preferences? (Prefix is always tulu3.4-sft-replica-50k-ultrafeedback-tpl THEN the model that judged it)