HGA-Thinker speech LM exported from SFT training (current train_sft.py format).
bridge.pt — HGA (s/b/c on 32 Whisper layers) + EMCA + audio boundary embeds
lora/ — PEFT LoRA adapter for the LLM (has_lora=True)
config.json — HGAThinkerConfig (llm_dim=3584)
processor_config.json — inference defaults
tokenizer.* — Qwen2.5 tokenizer
preprocessor_config.json — WhisperFeatureExtractor config
configuration_hga_thinker.py, modeling_hga_thinker.py… See the full description on the dataset page:
https://huggingface.co/datasets/VOLBEM/sft-6k.