This HF data repository contains the Bulgarian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
Machine-translated from yahma/alpaca-cleaned into Bulgarian.
This data is intended to be used for Bulgarian instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:… See the full description on the dataset page:
https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-bg.