Evaluation uses a fixed stride derived from eval_ratio=0.01
(roughly 1 out of 100 samples).
Length buckets by character count: 0–20, 20–40, 40–80, 80–120, 120–200,
200–400, 400+
Results (vad-macbert-mix/best)
en-zh_cn_vad_clean.csv
mse_mean=0.043734
mae_mean=0.149322
pearson_mean=0.7335
en-zh_cn_vad_long_clean.csv
mse_mean=0.031895
mae_mean=0.131320
pearson_mean=0.7565
Notes:
400+ bucket Pearson is unstable due to small sample size; interpret with care.
Limitations
Labels are derived from an English VAD teacher and transferred via parallel
alignment, so they reflect the teacher’s bias and may not match human Chinese
annotations.
Subtitle corpora include translation artifacts and formatting noise; cleaned
versions mitigate but do not fully remove this.
Extreme-length sentences are under-represented; performance on 400+ chars
is not reliable.