This dataset is built upon the Elise dataset and re-encoded using NVIDIA’s NeMo Audio Codec into nano audio tokens.
It is designed for fine-tuning multimodal LLMs and speech systems (TTS/ASR) that rely on codec-based audio token representations.
text: transcription of the utterance.
speaker: speaker identifier (string).
nano_layer_1 … nano_layer_4: tokenized audio representations from the NVIDIA NeMo Nano Codec… See the full description on the dataset page:
https://huggingface.co/datasets/nineninesix/elise-en-nano-codec-dataset.