Views
No views yet
google/gemma-3-12b-it, with the language
model extracted from the multimodal checkpoint. The multimodal wrapper
(vision tower + multi-modal projector) is removed and the model is
reconstructed as a pure Gemma3ForCausalLM (model_type: gemma3_text).language_model.* weights, strip the prefix to model.*vision_tower.* and multi_modal_projector.*text_config (48 layers, sliding_window=1024, 5:1 hybrid)