Views
No views yet
num_nextn_predict_layers = 1 and the appended model.layers.61.* MTP module, including enorm, hnorm, eh_proj, embed_tokens, and shared_head tensors.0..7, and layer 61 uses the same 8 routed experts as the main MoE layers.