Formatting is compliant with ChatML. "input" is the context and "output" is the expected model output to train on.
See the repo file generate_dataset.py for exactly this dataset was generated.
An opinionated and filtered mix of the following datasets:
argilla/ultrafeedback-binarized-preferences-cleaned
heegyu/glaive-function-calling-v2-formatted
berkeley-nest/Nectar
argilla/distilabel-math-preference-dpo… See the full description on the dataset page:
https://huggingface.co/datasets/andysalerno/rainbowfish-v1.