Nov 19, 2025: ⚙️ We release a new version of our model, which supports polyphonic pronunciation control and improves the performance of emotion, speaking style, and paralinguistic editing.
Nov 12, 2025: 📦 We release the optimized inference code and model weights of Step-Audio-EditX (HuggingFace; ModelScope) and Step-Audio-Tokenizer(HuggingFace; ModelScope)
Nov 06, 2025: 👋 We release the technical report of Step-Audio-EditX.
Introduction
We are open-sourcing Step-Audio-EditX, a powerful 3B-parameter LLM-based Reinforcement Learning audio model specialized in expressive and iterative audio editing. It excels at editing emotion, speaking style, and paralinguistics, and also features robust zero-shot text-to-speech (TTS) capabilities.
📑 Open-source Plan
Inference Code
Online demo (Gradio)
Step-Audio-Edit-Benchmark
Model Checkpoints
Step-Audio-Tokenizer
Step-Audio-EditX
Step-Audio-EditX-Int4
Training Code
GRPO training
SFT training
PPO training
⏳ Feature Support Plan
Editing
Polyphone pronunciation control
More paralinguistic tags ([Cough, Crying, Stress, etc.])
Speaking in a hearty, outgoing, and straight-talking manner
recite
Speaking in a clear, well-paced, poetry-reading manner
act_coy
Speaking in a sweet, playful, and endearing manner
warm
Speaking in a warm, friendly manner
shy
Speaking in a shy, timid manner
comfort
Speaking in a comforting, reassuring manner
authority
Speaking in an authoritative, commanding manner
chat
Speaking in a casual, conversational manner
radio
Speaking in a radio-broadcast manner
soulful
Speaking in a heartfelt, deeply emotional manner
gentle
Speaking in a gentle, soft manner
story
Speaking in a narrative, audiobook-style manner
vivid
Speaking in a lively, expressive manner
program
Speaking in a show-host/presenter manner
news
Speaking in a news broadcasting manner
advertising
Speaking in a polished, high-end commercial voiceover manner
roar
Speaking in a loud, deep, roaring manner
murmur
Speaking in a quiet, low manner
shout
Speaking in a loud, sharp, shouting manner
deeply
Speaking in a deep and low-pitched tone
loudly
Speaking in a loud and high-pitched tone
paralinguistic
[sigh]
Sighing sound
[inhale]
Inhaling sound
[laugh]
Laughter sound
[chuckle]
Chuckling sound
[exhale]
Exhaling sound
[clears throat]
Throat clearing sound
[snort]
Snorting sound
[giggle]
Giggling sound
[cough]
Coughing sound
[breath]
Breathing sound
[uhm]
Hesitation sound: "Uhm"
[Confirmation-en]
Confirming: "En"
[Surprise-oh]
Expressing surprise: "Oh"
[Surprise-ah]
Expressing surprise: "Ah"
[Surprise-wa]
Expressing surprise: "Wa"
[Surprise-yo]
Expressing surprise: "Yo"
[Dissatisfaction-hnn]
Dissatisfied sound: "Hnn"
[Question-ei]
Questioning: "Ei"
[Question-ah]
Questioning: "Ah"
[Question-en]
Questioning: "En"
[Question-yi]
Questioning: "Yi"
[Question-oh]
Questioning: "Oh"
Feature Requests & Wishlist
💡 We welcome all ideas for new features! If you'd like to see a feature added to the project, please start a discussion in our Discussions section.
We'll be collecting community feedback here and will incorporate popular suggestions into our future development plans. Thank you for your contribution!
Such legislation was clarified and extended from time to time thereafter. No, the man was not drunk, he wondered how we got tied up with this stranger. Suddenly, my reflexes had gone. It's healthier to cook without sugar.
For detailed quantization options and parameters, see quantization/README.md.
Technical Details
Step-Audio-EditX comprises three primary components:
A dual-codebook audio tokenizer, which converts reference or input audio into discrete tokens.
An audio LLM that generates dual-codebook token sequences.
An audio decoder, which converts the dual-codebook token sequences predicted by the audio LLM back into audio waveforms using a flow matching approach.
Audio-Edit enables iterative control over emotion and speaking style across all voices, leveraging large-margin data during SFT and PPO training.
Evaluation
Comparison between Step-Audio-EditX and Closed-Source models.
Step-Audio-EditX demonstrates superior performance over Minimax and Doubao in both zero-shot cloning and emotion control.
Emotion editing of Step-Audio-EditX significantly improves the emotion-controlled audio outputs of all three models after just one iteration. With further iterations, their overall performance continues to improve.
Generalization on Closed-Source Models.
For emotion and speaking style editing, the built-in voices of leading closed-source systems possess considerable in-context capabilities, allowing them to partially convey the emotions in the text. After a single editing round with Step-Audio-EditX, the emotion and style accuracy across all voice models exhibited significant improvement. Further enhancement was observed over the next two iterations, robustly demonstrating our model's strong generalization.
For paralinguistic editing, after editing with Step-Audio-EditX, the performance of paralinguistic reproduction is comparable to that achieved by the built-in voices of closed-source models when synthesizing native paralinguistic content directly. (sub means replacement of paralinguistic tags with native words)
Table: Generalization of Emotion, Speaking Style, and Paralinguistic Editing on Closed-Source Models.
Language
Model
Emotion ↑
Speaking Style ↑
Paralinguistic ↑
Iter0
Iter1
Iter2
Iter3
Iter0
Iter1
Iter2
Iter3
Iter0
sub
Iter1
Chinese
MiniMax-2.6-hd
71.6
78.6
81.2
83.4
36.7
58.8
63.1
67.3
1.73
2.80
2.90
Doubao-Seed-TTS-2.0
67.4
77.8
80.6
82.8
38.2
60.2
65.0
64.9
1.67
2.81
2.90
GPT-4o-mini-TTS
62.6
76.0
77.0
81.8
45.9
64.0
65.7
69.7
1.71
2.88
2.93
ElevenLabs-v2
60.4
74.6
77.4
79.2
43.8
63.3
69.7
70.8
1.70
2.71
2.92
English
MiniMax-2.6-hd
55.0
64.0
64.2
66.4
51.9
60.3
62.3
64.3
1.72
2.87
2.88
Doubao-Seed-TTS-2.0
53.8
65.8
65.8
66.2
47.0
62.0
62.7
62.3
1.72
2.75
2.92
GPT-4o-mini-TTS
56.8
61.4
64.8
65.2
52.3
62.3
62.4
63.4
1.90
2.90
2.88
ElevenLabs-v2
51.0
61.2
64.0
65.2
51.0
62.1
62.6
64.0
1.93
2.87
2.88
Average
MiniMax-2.6-hd
63.3
71.3
72.7
74.9
44.2
59.6
62.7
65.8
1.73
2.84
2.89
Doubao-Seed-TTS-2.0
60.6
71.8
73.2
74.5
42.6
61.1
63.9
63.6
1.70
2.78
2.91
GPT-4o-mini-TTS
59.7
68.7
70.9
73.5
49.1
63.2
64.1
66.6
1.81
2.89
2.90
ElevenLabs-v2
55.7
67.9
70.7
72.2
47.4
62.7
66.1
67.4
1.82
2.79
2.90
Acknowledgements
Part of the code and data for this project comes from:
Thank you to all the open-source projects for their contributions to this project!
License Agreement
The code in this open-source repository is licensed under the Apache 2.0 License.
Citation
@misc{yan2025stepaudioeditxtechnicalreport,
title={Step-Audio-EditX Technical Report},
author={Chao Yan and Boyong Wu and Peng Yang and Pengfei Tan and Guoqiang Hu and Yuxin Zhang and Xiangyu and Zhang and Fei Tian and Xuerui Yang and Xiangyu Zhang and Daxin Jiang and Gang Yu},
year={2025},
eprint={2511.03601},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2511.03601},
}
⚠️ Usage Disclaimer
Do not use this model for any unauthorized activities, including but not limited to:
Voice cloning without permission
Identity impersonation
Fraud
Deepfakes or any other illegal purposes
Ensure compliance with local laws and regulations, and adhere to ethical guidelines when using this model.
The model developers are not responsible for any misuse or abuse of this technology.
We advocate for responsible generative AI research and urge the community to uphold safety and ethical standards in AI development and application. If you have any concerns regarding the use of this model, please feel free to contact us.