Views
No views yet
| File | Stage | Size |
|---|---|---|
whisper_s2tt_ws2tt_lv3.pt | S2TT text stage: fine-tuned Whisper large-v3 encoder + custom FiLM decoder (whisper_s2tt.py) | 2.7 GB |
codec_decoder_l9_head.pt | Single-stream S2UT unit predictor: CodecDecoderXAR, n_q=1, mHuBERT-147 layer-9 units (codec_decoder_xattn.py) | 122 MB |
unit_vocoder_hifigan.pt | Dedicated Turkic single-stream unit vocoder (train_unit_voc_hifigan.py) | 338 MB |
.pt files are raw state_dict checkpoints matching the model classes
defined there (WhisperS2TT, CodecDecoderXAR, SpeechBrain UnitHifiganGenerator).RESULTS.md.demo_audio/ folder — 8 example system outputs (kaz source →
predicted tur/tat/uzb speech) from the final single-stream S2UT system.