This repository contains the pre-computed random unbalanced subsets used as baselines in the paper "A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't)".
Instruction fine-tuning of large language models (LLMs) often involves… See the full description on the dataset page:
https://huggingface.co/datasets/Harvard-DCML/tis-dolci-random-unbalanced.