This dataset contains two splits:
train — the RL/preference dataset from rl_data.jsonl with prompt, chosen, and rejected columns. This is the default split.
original — the original chat dataset from data.jsonl with a messages column.
Original Dataset