This dataset is an Uzbek translated version of OASST2 dataset.
Llama3 chat template + thread formatted dataset based on this translation is also available for model fine-tuning here.
The Uzbek translation was completed in 45 hours using a single T4 GPU and nllb-200-3.3B model.
Based on nllb metrics, you might want to only filter out records that were not originally in English or Russian since… See the full description on the dataset page:
https://huggingface.co/datasets/MLDataScientist/oasst2_uzbek.