xzuyn/lima-alpaca but outputs reworded to sound like Yoda. This currently contains all the rows that had an output of 1280 tokens or less. Nice, it is.
DatasetDict({
train: Dataset({
features: ['instruction', 'input', 'output', 'output_original'],
num_rows: 951
})
})
Original Dataset by Meta AI
LIMA: Less Is More Alignment