Based on Yodas2, this dataset contains interleaved (at the utterance level) text-audio documents derived from Yodas2-Mimi tokenized. It is designed for pretraining audio language models that can process both text and audio in a unified format.
The audio is represented as unicode strings (converted from Mimi audio codec tokens), making it compatible with standard language model training pipelines.… See the full description on the dataset page: https://huggingface.co/datasets/soda-research/yodas2-mm-pretrain.