This is an instruction dataset for fine-tuning in DPO.
The dataset consists of 981 training items and 33 test instances.
Each row in the dataset includes a column for facts, one for rules, another for positive examples of dialogue, as well as examples of dialogues to discard.
These components are concatenated to construct a prompt structure as follows:
Here is a synopsis of the bot's knowledge:
{memory}
The regulations are as follows:… See the full description on the dataset page:
https://huggingface.co/datasets/fractalego/wafl-functions-dataset.