Views
No views yet
transformers.models.moss_audio_tokenizer module. It is hosted as a Hugging Face Hub model repository and should be
loaded with trust_remote_code=True.
1import torch
2from transformers import AutoModel
3import torchaudio
4
5repo_id = "OpenMOSS-Team/MOSS-Audio-Tokenizer-v2"
6device = "cuda" if torch.cuda.is_available() else "cpu"
7model = AutoModel.from_pretrained(repo_id, trust_remote_code=True).eval().to(device)
8
9audio_path = "demo/demo_gt.wav" # replace with your own 48 kHz stereo audio path if needed
10wav, sr = torchaudio.load(audio_path)
11if sr != model.sampling_rate:
12 wav = torchaudio.functional.resample(wav, sr, model.sampling_rate)
13if wav.shape[0] == 1:
14 wav = wav.repeat(model.config.number_channels, 1)
15else:
16 wav = wav[: model.config.number_channels]
17wav = wav.unsqueeze(0).to(device)
18enc = model.encode(wav, return_dict=True)
19print(f"enc.audio_codes.shape: {enc.audio_codes.shape}")
20dec = model.decode(enc.audio_codes, return_dict=True)
21print(f"dec.audio.shape: {dec.audio.shape}")
22wav = dec.audio.squeeze(0)
23torchaudio.save("demo/demo_rec.wav", wav.cpu(), sample_rate=model.sampling_rate)
24
25# Decode using only the first 8 layers of the RVQ
26dec_rvq8 = model.decode(enc.audio_codes[:8], return_dict=True)
27wav_rvq8 = dec_rvq8.audio.squeeze(0)
28torchaudio.save("demo/demo_rec_rvq8.wav", wav_rvq8.cpu(), sample_rate=model.sampling_rate)trust_remote_code=True, pin revision to a reviewed commit hash.config.attention_implementation controls whether transformer layers prefer sdpa or flash_attention_2.
config.compute_dtype controls the non-quantizer autocast dtype and supports fp32, bf16.
config.codec_weight_dtype controls encoder/decoder parameter dtype and defaults to fp32.
The quantizer is always kept in fp32.1import torch
2from transformers import AutoModel
3
4repo_id = "OpenMOSS-Team/MOSS-Audio-Tokenizer-v2"
5device = "cuda" if torch.cuda.is_available() else "cpu"
6model = AutoModel.from_pretrained(repo_id, trust_remote_code=True, low_cpu_mem_usage=True, codec_weight_dtype="bf16").eval().to(device)1model.set_attention_implementation("flash_attention_2")
2model.set_compute_dtype("bf16")
3model.set_codec_weight_dtype("bf16") # encoder/decoder bf16, quantizer fp32model.to(torch.bfloat16) on the whole codec; that also casts quantizer weights and can cause dtype mismatches or serious precision loss.MossAudioTokenizerModel.encode, decode, batch_encode, and batch_decode all support streaming through a
chunk_duration argument.chunk_duration is expressed in seconds.chunk_duration * MossAudioTokenizerConfig.sampling_rate must be divisible by MossAudioTokenizerConfig.downsample_rate.(2, T) or batched stereo inputs shaped (B, 2, T).1import torch
2from transformers import AutoModel
3
4repo_id = "OpenMOSS-Team/MOSS-Audio-Tokenizer-v2"
5device = "cuda" if torch.cuda.is_available() else "cpu"
6model = AutoModel.from_pretrained(repo_id, trust_remote_code=True).eval().to(device)
7audio = torch.randn(2, 48000 * 6).to(device) # dummy stereo waveform
8
9# 6.0s @ 48kHz = 288000 samples, divisible by downsample_rate=3840
10enc = model.encode(audio.unsqueeze(0), return_dict=True, chunk_duration=0.08)
11dec = model.decode(enc.audio_codes, return_dict=True, chunk_duration=0.08)
12
13batch_enc = model.batch_encode([audio, audio[:, : 48000 * 3]], chunk_duration=0.08)
14codes_list = [
15 batch_enc.audio_codes[:, i, : batch_enc.audio_codes_lengths[i]]
16 for i in range(batch_enc.audio_codes.shape[1])
17]
18batch_dec = model.batch_decode(codes_list, chunk_duration=0.08)configuration_moss_audio_tokenizer.pymodeling_moss_audio_tokenizer.py__init__.pyconfig.jsonmodel.safetensors.index.jsonmodel-00001-of-00003.safetensors, model-00002-of-00003.safetensors,
model-00003-of-00003.safetensorsdemo/demo_gt.wav| Model | bps | Frame rate | Nq | Speech: SIM ↑ (EN/ZH) | Speech: STOI ↑ (EN/ZH) | Speech: PESQ-NB ↑ (EN/ZH) | Speech: PESQ-WB ↑ (EN/ZH) | Audio/Music: Mel-Loss ↓ | Audio/Music: STFT-Dist. ↓ |
|---|---|---|---|---|---|---|---|---|---|
| XCodec2.0 | 800 | 50 | 1 | 0.82 / 0.74 | 0.92 / 0.86 | 3.04 / 2.46 | 2.43 / 1.96 | -- / -- | -- / -- |
| MiMo Audio Tokenizer | 850 | 25 | 4 | 0.80 / 0.74 | 0.91 / 0.87 | 2.94 / 2.62 | 2.39 / 2.14 | 0.82 / 0.81 | 2.33 / 2.23 |
| Higgs Audio Tokenizer | 1000 | 25 | 4 | 0.77 / 0.68 | 0.83 / 0.82 | 3.03 / 2.61 | 2.48 / 2.14 | 0.83 / 0.80 | 2.20 / 2.05 |
| SpeechTokenizer | 1000 | 50 | 2 | 0.36 / 0.25 | 0.77 / 0.68 | 1.59 / 1.38 | 1.25 / 1.17 | -- / -- | -- / -- |
| XY-Tokenizer | 1000 | 12.5 | 8 | 0.85 / 0.79 | 0.92 / 0.87 | 3.10 / 2.63 | 2.50 / 2.12 | -- / -- | -- / -- |
| BigCodec | 1040 | 80 | 1 | 0.84 / 0.69 | 0.93 / 0.88 | 3.27 / 2.55 | 2.68 / 2.06 | -- / -- | -- / -- |
| Mimi | 1100 | 12.5 | 8 | 0.74 / 0.59 | 0.91 / 0.85 | 2.80 / 2.24 | 2.25 / 1.78 | 1.24 / 1.19 | 2.62 / 2.49 |
| MOSS-Audio-Tokenizer-v2 (Ours) | 750 | 12.5 | 6 | 0.82 / 0.75 | 0.92 / 0.88 | 3.14 / 2.68 | 2.59 / 2.19 | 0.93 / 0.91 | 2.28 / 2.14 |
| MOSS-Audio-Tokenizer-v2 (Ours) | 1000 | 12.5 | 8 | 0.88 / 0.80 | 0.94 / 0.90 | 3.39 / 2.93 | 2.88 / 2.43 | 0.88 / 0.86 | 2.22 / 2.07 |
| — | — | — | — | — | — | — | — | — | — |
| DAC | 1500 | 75 | 2 | 0.48 / 0.41 | 0.83 / 0.79 | 1.87 / 1.67 | 1.48 / 1.37 | -- / -- | -- / -- |
| Encodec | 1500 | 75 | 2 | 0.60 / 0.45 | 0.85 / 0.81 | 1.94 / 1.80 | 1.56 / 1.48 | 1.12 / 1.04 | 2.60 / 2.42 |
| Higgs Audio Tokenizer | 2000 | 25 | 8 | 0.90 / 0.83 | 0.85 / 0.85 | 3.59 / 3.22 | 3.11 / 2.73 | 0.74 / 0.70 | 2.07 / 1.92 |
| SpeechTokenizer | 2000 | 50 | 4 | 0.66 / 0.50 | 0.88 / 0.80 | 2.38 / 1.79 | 1.92 / 1.49 | -- / -- | -- / -- |
| Qwen3 TTS Tokenizer | 2200 | 12.5 | 16 | 0.95 / 0.88 | 0.96 / 0.93 | 3.66 / 3.10 | 3.19 / 2.62 | -- / -- | -- / -- |
| MiMo Audio Tokenizer | 2250 | 25 | 12 | 0.89 / 0.83 | 0.95 / 0.92 | 3.57 / 3.25 | 3.05 / 2.71 | 0.70 / 0.68 | 2.21 / 2.10 |
| Mimi | 2475 | 12.5 | 18 | 0.89 / 0.76 | 0.94 / 0.91 | 3.49 / 2.90 | 2.97 / 2.35 | 1.10 / 1.06 | 2.45 / 2.32 |
| MOSS-Audio-Tokenizer-v2 (Ours) | 1500 | 12.5 | 12 | 0.93 / 0.86 | 0.95 / 0.92 | 3.66 / 3.24 | 3.23 / 2.77 | 0.83 / 0.79 | 2.15 / 1.98 |
| MOSS-Audio-Tokenizer-v2 (Ours) | 2000 | 12.5 | 16 | 0.95 / 0.89 | 0.96 / 0.94 | 3.80 / 3.44 | 3.45 / 3.01 | 0.79 / 0.75 | 2.10 / 1.93 |
| — | — | — | — | — | — | — | — | — | — |
| DAC | 3000 | 75 | 4 | 0.74 / 0.67 | 0.90 / 0.88 | 2.76 / 2.47 | 2.31 / 2.07 | 0.86 / 0.83 | 2.23 / 2.10 |
| MiMo Audio Tokenizer | 3650 | 25 | 20 | 0.91 / 0.85 | 0.95 / 0.93 | 3.73 / 3.44 | 3.25 / 2.89 | 0.66 / 0.65 | 2.17 / 2.06 |
| SpeechTokenizer | 4000 | 50 | 8 | 0.85 / 0.69 | 0.92 / 0.85 | 3.05 / 2.20 | 2.60 / 1.87 | -- / -- | -- / -- |
| Mimi | 4400 | 12.5 | 32 | 0.94 / 0.83 | 0.96 / 0.94 | 3.80 / 3.31 | 3.43 / 2.78 | 1.02 / 0.98 | 2.34 / 2.21 |
| Encodec | 4500 | 75 | 6 | 0.86 / 0.75 | 0.92 / 0.91 | 2.91 / 2.63 | 2.46 / 2.15 | 0.91 / 0.84 | 2.33 / 2.17 |
| DAC | 6000 | 75 | 8 | 0.89 / 0.84 | 0.95 / 0.94 | 3.75 / 3.57 | 3.41 / 3.20 | 0.65 / 0.63 | 1.97 / 1.87 |
| MOSS-Audio-Tokenizer-v2 (Ours) | 3000 | 12.5 | 24 | 0.96 / 0.92 | 0.97 / 0.95 | 3.94 / 3.64 | 3.66 / 3.28 | 0.75 / 0.71 | 2.04 / 1.87 |
| MOSS-Audio-Tokenizer-v2 (Ours) | 4000 | 12.5 | 32 | 0.97 / 0.93 | 0.97 / 0.96 | 3.98 / 3.72 | 3.75 / 3.39 | 0.73 / 0.69 | 2.02 / 1.84 |
![]() |
1@misc{gong2026mossaudiotokenizerscaling,
2 title={MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models},
3 author={Yitian Gong and Kuangwei Chen and Zhaoye Fei and Xiaogui Yang and Ke Chen and Yang Wang and Kexin Huang and Mingshu Chen and Ruixiao Li and Qingyuan Cheng and Shimin Li and Xipeng Qiu},
4 year={2026},
5 eprint={2602.10934},
6 archivePrefix={arXiv},
7 primaryClass={cs.SD},
8 url={https://arxiv.org/abs/2602.10934}
9}LICENSE for the full license text.