This is the dataset used for creating aslon1213orpheus-qa-dataset.
This dataset is used by other notebook to create aslon1213orpheus-qa-dataset and embed tokenized versions of the text.
The idea is to create a QA pair dataset to finetune the Base LLM model for more contextual appropriate TTS generation.
For Question part uzbek text corpus was used and Answer part was generated using grok-4.0-fast model.