Views
No views yet




<think> and </think> tags. The instruct-mode prefix is task-specific. Prepend one of the following to the assistant's response, i.e. right after the assistant ChatML header:| Task | Mode | Prefix | Note |
|---|---|---|---|
| Text | thinking | <think>\n | The model closes the trace with </think> and then answers. |
| Text | instruct (non-thinking) | </think>\n | |
| Audio | instruct (non-thinking) | <think></think> | No newline between </think> and the answer. |
mamba-ssm and causal-conv1d. Build against your CUDA toolchain:1python3 -m pip install transformers==5.14.0 safetensors==0.8.0
2python3 -m pip install --no-build-isolation causal-conv1d==1.6.2.post1 mamba-ssm==2.3.1vllm/vllm-openai:v0.20.0 image does not include audio codecs. This command installs audio-related packages: python3 -m pip install "vllm[audio]".LLM.generate and an OpenAI-compatible audio_url server.<sound>\n placeholder:1[
2 {
3 "id": "sample_0",
4 "sound": "/path/to/audio_0.wav",
5 "conversations": [
6 {"from": "human", "value": "<sound>\nDescribe the audio in detail."},
7 {"from": "gpt", "value": "N/A"}
8 ]
9 },
10 {
11 "id": "sample_1",
12 "sound": "/path/to/audio_1.wav",
13 "conversations": [
14 {"from": "human", "value": "<sound>\n{prompt}"},
15 {"from": "gpt", "value": "N/A"}
16 ]
17 },
18 ...
19]1python3 -m pip install "vllm[audio]" # audio input decoding; skip if your image already bundles it (see Environment)
2pip install -e inference_scripts_vllm/audioqa_scripts --no-deps --no-build-isolation1python inference_scripts_vllm/audioqa_scripts/run_audioqa_vllm.py \
2 --model-path "$(pwd)/checkpoint_folder_full" \
3 --input-json ./inputs.json \
4 --output-jsonl ./audioqa_outputs/results.jsonl \
5 --tensor-parallel-size 81bash inference_scripts_vllm/audioqa_scripts/serve_audioqa_vllm.sh "$(pwd)/checkpoint_folder_full" 8000
2python inference_scripts_vllm/audioqa_scripts/client_audioqa.py --audio /path/to/audio.wav --prompt "Describe this audio."bash inference_scripts_hf/inference_example.sh.Transcribe the speech in the input audio.\n<sound>; speech translation — a translation instruction such as Translate the speech in the input audio into English.\n<sound>.bash model_conversion_scripts/prepare_audiogen_vllm_checkpoint.sh (which only creates symlinks of safetensors under checkpoint_folder_audiogen).hf download hf-audio/xcodec-hubert-general-balanced --local-dir /path/to/xcodec1/path/to/caption_txt_dir/ with all .txt files where each contains one caption. Run cd inference_scripts_vllm/audiogen_scripts/ and
set --tensor-parallel-size to the number of GPUs. RunXCODEC1_PATH=/path/to/xcodec1 python3 run_audio_gen_vllm_rvq_logit_mask.py \
--task tta \
--model-path $(pwd)/../../checkpoint_folder_audiogen/ \
--dataset-path /path/to/caption_txt_dir/ \
--output-dir ../../tta_outputs/dataset_name/ \
--tensor-parallel-size 8 \
--temperature 1.0 \
--top-k 80 \
--max-tokens 2048 \
--cfg-scale 3.0 \
--cfg-pairs-per-batch 2enhancement_VAE/README.md.audex_causal_speech_decoder (default)../run_tts_vllm.sh --transcription "The weather is so good, and I want to enjoy the beautiful morning in the park." \
--output-dir ./tts_outputs --utt-id the_weather_is_so_good<think>\n prefix. Text instruct (non-thinking) mode uses </think>\n. See run_text_vllm_example.py which implements both (--disable-thinking selects instruct mode).chat_template.jinja uses <think></think> when enable_thinking=false which is used for audio tasks. To reproduce the reported text instruct-mode results, construct the prompt directly with </think>\n instead of relying on apply_chat_template.python model_conversion_scripts/convert_full_HF_to_textonly_HF.py to remove the audio-related vocabularies.cd inference_scripts_vllm/textonly_scripts/; python run_text_vllm_example.py --model-path $(pwd)/../../checkpoint_folder_textonly.sampling_params = SamplingParams(allowed_token_ids=list(range(131072))) in vLLM inference to mask the audio tokens, although we did not thoroughly test this approach.inference_scripts_vllm/unified_s2s_scripts/README.md.







@article{Nemotron-Labs-Audex,
title={Unified Audio Intelligence Without Regressing on Text Intelligence},
author={Kong, Zhifeng and Lee, Sang-gil and Kim, Jaehyeon and Wang, Boxin and Liu, Zihan and Kim, Sungwon and Chen, Yang and Goel, Arushi and Roy, Rajarshi and Dai, Wenliang and Yang, Zhuolin and Chen, Yangyi and Jiang, Dongfu and Ghosh, Sreyan and Rintamaki, Tuomas and Tao, Andrew and Raiman, Jonathan and Shoeybi, Mohammad and Catanzaro, Bryan and Ping, Wei},
year={2026}
}