Voice Chat Pipeline -> ASR + TurnDetector + VAD + LLM + TTS
In the Voice Chat Pipeline, if we only rely on VAD (Voice Activity Detection) to determine whether the user's current turn input has ended, we cannot accurately handle situations where users pause while thinking. When there are pauses during the current turn input that hasn't been completed yet, VAD will detect the pause and prematurely judge that the sentence has ended, but semantically the sentence is not yet complete.
This introduces the Turn-Detector Model. The turn detection model is mainly applied in voice + text modal dialogue scenarios. At the semantic level, the turn detection model can analyze the text information transcribed by the ASR model at the semantic level, more accurately determining whether the current user input has ended. The Turn-Detector Model chooses small-parameter (0.5B/0.6B) large models based on Transformer architecture that have undergone instruction fine-tuning, with the main task being to predict the probability of the next_token being <|im_end|>.
Task: Semantic-level turn recognition, predicting the probability of next_token being <|im_end|>
Model: Small-parameter models after instruction fine-tuning (Qwen2.5-0.5B-Instruct, Qwen3-0.6B)
Goal: Reduce inaccurate VAD interruptions in voice dialogue pipelines (e.g., pauses caused while thinking of the next word)
# 1. get the user input
How tall is the Eiffel Tower
# 2. apply_chat_template
<|im_start|>user<|im_sep|>How tall is the Eiffel Tower<|im_end|>
# 3. cut <|im_end|>
<|im_start|>user<|im_sep|>How tall is the Eiffel Tower
# 4. predict next token
The turn detection model is mainly applied in Chinese and English voice + text modal dialogue scenarios, with input data types mostly being common text instruction data and colloquial chat dialogue data. Therefore, the dataset uses public datasets such as Alpaca, MagicData (ASR dialogue dataset), ShareChatX, etc.
Alpaca
Magicdata
ShareChatX
Characteristics of ASR transcribed text
Sometimes sentence endings don't contain punctuation marks
There may be filler words or ... during the process
Dataset optimization based on ASR transcribed text characteristics
Sentence filtering: Call large models to analyze current input content, retaining semantically complete and colloquial data from the dataset
Filler word insertion: Randomly insert 1 filler word in sentences to simulate the actual effect of spoken dialogue
Call large models to generate Chinese and English filler word tables
1[2{3"instruction":"How tall is the Eiffel Tower",4"input":"",5"output":""6},7{8"instruction":"How tall is the um... Eiffel Tower",9"input":"",10"output":""11},12{13"instruction":"Um how tall is the Eiffel Tower",14"input":"",15"output":""16},17 ...
18]