MOSS-Audio-Tokenizer-v2 codes (12-codebook RVQ @ 48 kHz) with captions for
scientifi-papers/a1-plus, ready for MOSS-TTS-Local-Transformer (moss 1.5 local)
training. No MP3s — tokens + text + captions only.
tokenized_audio.parquet, per row (key joins to a1-plus):
target_codes : int16 bytes, reshape [target_frames, 12]
ref_codes : int16 bytes, reshape [ref_frames, 12] (same-speaker reference; may be… See the full description on the dataset page:
https://huggingface.co/datasets/TTS-AGI/additional-data-a1-plus.