Views
No views yet
safetensors conversion of NVIDIA's NeMo NanoCodec
(nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps),
the 22.05 kHz / 21.5 fps / 1.89 kbps variant. It is a neural audio codec using finite scalar
quantization (GroupFSQ, 8 groups × FSQ levels [8, 7, 6, 6], 32 embed dims) with a HiFi-GAN
encoder/decoder..nemo archive that requires the NeMo/PyTorch stack (Linux-only in
practice). The tensors here are extracted from that archive verbatim (weight-norm kept in the
modern parametrizations.weight.{original0,original1} layout; no numerics changed). It is the
codec layer of the Gepard-1.0 streaming-TTS port
(xocialize/mlx-gepard-swift).| file | what |
|---|---|
model.safetensors | encoder (audio_encoder.*) + decoder (audio_decoder.*) + quantizer weights (fp32) |
codec_config.yaml | the NeMo config the weights were exported with (rates, dims, FSQ levels) |
nvidia-open-model-license-agreement-june-2024.pdf | the governing license (shipped per §3.1) |
NOTICE | the §3.1 attribution notice |
NOTICE. The model is commercially usable; redistribution of the model and derivatives
is permitted with the §3.1 obligations (ship a copy of the Agreement and the Notice below).Licensed by NVIDIA Corporation under the NVIDIA Open Model License