This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (english).
sequence: Full LLASA training sequence with phonemes and XCodec2 tokens
transcription_full: Transcript matching the actual audio (left + right portions)
transcription_original: Original full transcript
removed_words: Words that were removed for infilling training
phonemes_annotated: Phoneme tokens with <|ph_space|>… See the full description on the dataset page:
https://huggingface.co/datasets/zuhri025/AE_english_data_stage_3_v2.