This is a full fine-tuned checkpoint of Qwen/Qwen3-ASR-1.7B for Japanese galgame, visual novel, and anime-style speech recognition.
The model was fine-tuned on litagin/Galgame_Speech_ASR_16kHz, a large Japanese galgame speech ASR corpus. This repository includes both inference weights and training recovery files so that others can resume training or continue domain adaptation.
The current checkpoint is intended as a domain-tuned Japanese ASR model. A small external benchmark is included below, comparing this model against the upstream Qwen3-ASR base models on both game/anime-style audio and compact general Japanese ASR sanity sets.
Intended Use
This model is intended for Japanese ASR in speech with galgame or anime-like delivery, including:
visual novel and game voice transcription
subtitle generation workflows
Japanese character dialogue with expressive voice acting
research on domain adaptation from general ASR models to anime-style speech
It is not yet validated as a general-purpose Japanese ASR model. For broad Japanese speech, compare against the original base model before production use.
Metric: strict character error rate (CER) after removing whitespace and common Japanese/ASCII punctuation. S, I, and D are substitution, insertion, and deletion rates divided by reference characters. The same decoding and normalization were used for all models.
model
rows
CER
S
I
D
Qwen/Qwen3-ASR-0.6B
800
0.1673
0.1025
0.0214
0.0434
jaykwok/Qwen3-ASR-0.6B-JA-Anime-Galgame
800
0.1438
0.0962
0.0228
0.0249
Qwen/Qwen3-ASR-1.7B
800
0.1437
0.0851
0.0169
0.0418
jaykwok/Qwen3-ASR-1.7B-JA-Anime-Galgame
800
0.1285
0.0812
0.0231
0.0242
CER by source:
model
Nekopara
Anime Speech
JSUT
Common Voice
Qwen/Qwen3-ASR-0.6B
0.2900
0.1244
0.1297
0.1552
jaykwok/Qwen3-ASR-0.6B-JA-Anime-Galgame
0.2392
0.0811
0.1207
0.1568
Qwen/Qwen3-ASR-1.7B
0.2803
0.1091
0.0948
0.1269
jaykwok/Qwen3-ASR-1.7B-JA-Anime-Galgame
0.2276
0.0799
0.0998
0.1312
For this 1.7B checkpoint, full SFT improves overall CER from 0.1437 to 0.1285, a 10.6% relative reduction. The largest improvement is deletion reduction, from 0.0418 to 0.0242. In-domain gains are stronger: Nekopara CER improves by 18.8% relative, and Anime Speech CER improves by 26.8% relative. JSUT and Common Voice are slightly worse than the 1.7B base in this small sample, so this checkpoint should still be treated primarily as a galgame/anime-domain model rather than a general Japanese ASR upgrade.
These numbers are a small reproducible sanity benchmark, not a comprehensive public leaderboard. Strict character CER can over-penalize kana/kanji variants, long-vowel spelling, expressive writing, and transcript style differences.
Additional Evaluation Candidates
Recommended additional evaluation sets:
ntaquan0125/steinsgate-voice, a relatively small STEINS;GATE visual novel voice dataset with Japanese audio and text, if access and licensing are acceptable
This fine-tuned checkpoint was trained on litagin/Galgame_Speech_ASR_16kHz. Users must review and comply with the dataset license and upstream terms before redistribution, commercial use, or further fine-tuning. This model card does not grant rights beyond the upstream model and dataset licenses.