We're excited to introduce Nemotron-Labs-Audex-2B, a unified audio-text LLM with a similar recipe as Nemotron-Labs-Audex-30B-A3B. Audex-2B extends the vocabulary for discrete audio tokens used for speech and general audio outputs, as well as an audio encoder for speech and general audio inputs. Audex-2B delivers strong abilities on audio tasks (audio understanding, speech recognition
and translation, text-to-speech, audio generation, and speech-to-speech generation) while preserving
very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM
backbone with marginal or no regression. Audex-2B operates in both thinking and instruct (non-thinking) modes.
Model Architecture
architecture
Templates
template
Multi-Stage-SFT and Cascaded-RL Pipelines
training
Audex-2B Details
Audex-2B is a compact, dense model trained using the same multi-stage SFT recipe as Audex-30B-A3B. Its smaller size reduces inference resource requirements. Audex-2B is released after multi-stage SFT, while Audex-30B-A3B additionally undergoes cascaded RL.
Quick Start
Audex-2B follows the ChatML template and supports both thinking and instruct (non-thinking) modes. Reasoning content is enclosed within <think> and </think> tags. To activate the instruct (non-thinking) mode, we prepend <think></think> to the beginning of the assistant’s response.
Audex-2B supports up to a 128K-token context length.
Audex-2B follows Nemotron-Cascade-2 on text evaluation.
Audex-2B has different recommended inference setups per audio-related task as described below.
vLLM inference — text-only reasoning, text-to-speech, text-to-audio, and audio understanding / speech recognition / speech translation: runs on vLLM 0.20.0.
Hugging Face / transformers inference — requires transformers >= 4.53.0 (tested with 4.53.3) and also works with transformers >= 5.0.
Audio extras:vllm/vllm-openai:v0.20.0 image does not include audio codecs. This command installs audio-related packages: python3 -m pip install "vllm[audio]".
vLLM plugin: Audex-2B is served through a small vLLM plugin. From the model root, register it once (needed for every vLLM task below): pip install -e nemotron_dense_vllm_plugin --no-deps --no-build-isolation.
Audio QA Inference
Audio QA includes audio understanding, speech recognition, and speech translation (see templates in Introduction).
vLLM (recommended) — offline LLM.generate and an OpenAI-compatible audio_url server.
Hugging Face / transformers — requires transformers >= 4.53.0 (we tested with 4.53.3), and also works with transformers >= 5.
Inputs
To prepare inputs, create a JSON file in the following format with the <sound>\n placeholder:
json
1[2{3"id":"sample_0",4"sound":"/path/to/audio_0.wav",5"conversations":[6{"from":"human","value":"<sound>\nDescribe the audio in detail."},7{"from":"gpt","value":"N/A"}8]9},10{11"id":"sample_1",12"sound":"/path/to/audio_1.wav",13"conversations":[14{"from":"human","value":"<sound>\n{prompt}"},15{"from":"gpt","value":"N/A"}16]17},18 ...
19]
Inference recipes
For audio understanding, we use top_p=0.9 and temperature=0.7.
For speech recognition and translation, we use greedy sampling.
vLLM (recommended)
To install environments:
bash
1python3 -m pip install"vllm[audio]"# audio input decoding; skip if your image already bundles it (see Environment)2pip install -e inference_scripts_vllm/audioqa_scripts --no-deps --no-build-isolation
1bash inference_scripts_vllm/audioqa_scripts/serve_audioqa_vllm.sh "$(pwd)/checkpoint_folder_full"80002python inference_scripts_vllm/audioqa_scripts/client_audioqa.py --audio /path/to/audio.wav --prompt "Describe this audio."
Hugging Face / transformers
Example inference script: bash inference_scripts_hf/inference_example.sh.
Task instruction examples: audio understanding — a question about the audio; speech recognition — Transcribe the speech in the input audio.\n<sound>; speech translation — a translation instruction such as Translate the speech in the input audio into English.\n<sound>.
Audio Generation Inference
Audio generation includes text-to-speech and text-to-audio generation.
First, prepare vLLM inference using bash model_conversion_scripts/prepare_audiogen_vllm_checkpoint.sh (which only creates symlinks of safetensors under checkpoint_folder_audiogen).
Prepare a folder /path/to/caption_txt_dir/ with all .txt files where each contains one caption. Run cd inference_scripts_vllm/audiogen_scripts/ and
set --tensor-parallel-size to the number of GPUs. Run
Finally (optional), apply the 48 kHz enhancement VAE to the generated waveforms; see enhancement_VAE/README.md.
Text-to-speech (TTS)
we recommend using the standalone Audex causal speech decoder in audex_causal_speech_decoder (default).
./run_tts_vllm.sh --transcription "The weather is so good, and I want to enjoy the beautiful morning in the park." \
--output-dir ./tts_outputs --utt-id the_weather_is_so_good
Alternatively, users can download the original XCodec2 via this repo and decode the tokens after full generation; this has better quality but is not streaming.
To create the checkpoint that completely matches their setup, run python model_conversion_scripts/convert_full_HF_to_textonly_HF.py to remove the audio-related vocabularies.
Then, the text-only inference completely follows Nemotron-Cascade-2-30B-A3B. See a simple reasoning example with cd inference_scripts_vllm/textonly_scripts/; python run_text_vllm_example.py --model-path $(pwd)/../../checkpoint_folder_textonly.
Note: you could also avoid model conversion by using sampling_params = SamplingParams(allowed_token_ids=list(range(131072))) in vLLM inference to mask the audio tokens, although we did not thoroughly test this approach.
Demo: speech-to-speech interaction with text reasoning
See inference_scripts_vllm/unified_s2s_scripts/README.md.
Reproducibility
The benchmark numbers below use the following setups:
Text-to-speech (TTS): the original non-streaming XCodec2 decoder.
Text-to-audio (TTA): XCodec1 followed by the enhancement VAE.
Text: same as Nemotron-Cascade-2.
Audio understanding: transformers 4.53.3 and Megatron-LM's native inference.
@article{Nemotron-Labs-Audex,
title={Unified Audio Intelligence Without Regressing on Text Intelligence},
author={Kong, Zhifeng and Lee, Sang-gil and Kim, Jaehyeon and Wang, Boxin and Liu, Zihan and Kim, Sungwon and Chen, Yang and Goel, Arushi and Roy, Rajarshi and Dai, Wenliang and Yang, Zhuolin and Chen, Yangyi and Jiang, Dongfu and Ghosh, Sreyan and Rintamaki, Tuomas and Tao, Andrew and Raiman, Jonathan and Shoeybi, Mohammad and Catanzaro, Bryan and Ping, Wei},
year={2026}
}