Views
No views yet
GrugMoeForCausalLM (model_type: grug_moe);
serve with the marin vLLM fork (--enable-expert-parallel), not upstream vLLM.step-42150 (July 2T cooldown).wildchat_386k (math-weak chat, no thinking), 1 packed epoch
= 257 steps, seq_len 32768, global batch 64, optimizer AdamH (fp32),
cut cross-entropy (batched_xla). Final training loss 0.676.marin-community/marin-tokenizer; chat template delphi_v0.jinja2.GrugModelConfig.hf_checkpoint_converter().with_config_overrides({"dtype":"bfloat16"})
path (experiments/grug/moe/model.py), reproducing
tests/vllm/e2e/test_june_67b_a2b_hf_bf16_export.py. pending_qb_betas is baked
into the router bias before export (required for correct logits). All tensors BF16.s3://marin-us-east-02a/marin/exports/grug/june-67b-a2b-sft-s1-wildchat/step-257/hf-bf16-vllm/