Belarusian OpenOrca dataset - is rich collection of augmented FLAN data aligns, that translated in belarusian language.
That dataset should help training LLM in belarusian language and should help on other NLP tasks.
This dataset have 2 version:
~1M GPT-4 completions (Now translating)
~3.2M GPT-3.5 completions (Can be translated in future)
'id', a unique numbered identifier which includes one of 'niv'⦠See the full description on the dataset page:
https://huggingface.co/datasets/WiNE-iNEFF/1M-OpenOrca_be.