This dataset is a collection of prompted examples from P3, NI, RST, BigBench, FLAN & StackExchange,
and examples from C4. The C4 examples are labeled "not-task-like" and the P3, NI, RST, BigBench, FLAN,
StackExchange & UnNatural Instructions examples are "task-like". Examples were sampled from C4 so that
the distribution of example lengths is similar for C4, and P3, NI, RST, BigBench, FLAN, StackExchange
& UnNatural Instructions examples. Some datasets from P3 were ignored because their examples were too
long. Some datasets from P3, BigBench, FLAN, StackExchange & UnNatural Instructions are held out for
validation. The datasets from the train split of Natural Instuctions were used for creating the train
set of the tasky data while those from the test split were used in creating the validation set.
Non-tasky validation data was gathered from C4 without intentionally matching the length distribution.
Tasky validation data was gathered from the validation set of certain held-out datasets from P3, NI,
BigBench, FLAN, StackExchange & UnNatural Instructions.