Supervised Fine-Tuning (SFT) dataset for OGBert, containing prompt-completion pairs generated from the OpenGloss dictionary.
This dataset transforms dictionary entries into instruction-following format for fine-tuning language models. Each entry generates multiple training examples using different lexical properties:
Definitions: "Definition: X\nWhich word is defined?" → word
Synonyms: "Synonyms: X, Y, Z\nWhich word fits these… See the full description on the dataset page:
https://huggingface.co/datasets/mjbommar/ogbert-v1-sft.