Views
No views yet
Fork: transformers==5.13.0 compatibility
This is a fork ofmoonshotai/Moonlight-16B-A3B-Instructwhose only differences from upstream are edits to the custommodeling_deepseek.pyfile so that the model loads and generates correctly ontransformers==5.13.0. The safetensors weights are identical to upstream. Verified working withtorch==2.6.0+cu124andtransformers==5.13.0.The upstream file targetstransformers==4.48.2. Between 4.48 and 5.13 several internal APIs changed or were removed, and the transformers-5.x meta-device-first load path caused an uninitialized RoPE buffer that silently broke the model. Concretely, the following was changed inmodeling_deepseek.py:
- RoPE
inv_freqbuffer (the critical fix). In transformers ≥ 5,from_pretrainedinitializes modules on themetadevice before loading the checkpoint. Buffers registered withpersistent=False(likeDeepseekV3RotaryEmbedding.inv_freq) are not populated fromstate_dict, soinv_freqremained as meta-device zeros — every position producedcos=1, sin=0and RoPE was silently disabled.DeepseekV3RotaryEmbeddingnow lazily recomputesinv_freqon the real device inside_set_cos_sin_cacheif it is onmetaor all-zero. Without this fix the model produces fluent but position-blind gibberish.rope_scalingdefault schema. transformers ≥ 4.45 auto-populatesconfig.rope_scalingwith{'rope_type': 'default', ...}even when the config file has norope_scalingkey._init_rope(and the softmax-scale branch) now call a new_rope_scaling_active()helper that treatsrope_type == 'default'(or a missingfactor) as "no scaling," matching the pre-5.x behavior for Moonlight.Cache.seen_tokensremoved.prepare_inputs_for_generationnow usesgetattr(past_key_values, "seen_tokens", cache_length)wherecache_lengthcomes fromget_seq_length().Cache.get_usable_length(...)removed. All call sites were replaced with the equivalentget_seq_length(...)(per-layer where applicable).DynamicCachehas no cache size limit, so the replacement is exact.DynamicCache.get_max_cache_shape()changed return value. It now returns-1(previouslyNone) for unbounded caches.prepare_inputs_for_generationtreats any non-positive value as "unbounded," which prevents an erroneousattention_mask[:, -(-1):]truncation that dropped the first token.- Full-sequence
position_idsin the decode loop. The new_sampleloop passes position IDs covering the whole sequence even when only one new token is being processed.prepare_inputs_for_generationnow trimsposition_idsdown toinput_ids' length so RoPE indexing stays consistent.The README's two "Inference with Hugging Face Transformers" snippets below have been updated for the new API (the second snippet usesreturn_dict=Truewithapply_chat_template, which is required in transformers ≥ 5) and both were re-run end-to-end against this fork — the actual outputs are shown after each snippet.

| Benchmark (Metric) | Llama3.2-3B | Qwen2.5-3B | DSV2-Lite | Moonlight | |
|---|---|---|---|---|---|
| Activated Param† | 2.81B | 2.77B | 2.24B | 2.24B | |
| Total Params† | 2.81B | 2.77B | 15.29B | 15.29B | |
| Training Tokens | 9T | 18T | 5.7T | 5.7T | |
| Optimizer | AdamW | * | AdamW | Muon | |
| English | MMLU | 54.75 | 65.6 | 58.3 | 70.0 |
| MMLU-pro | 25.0 | 34.6 | 25.5 | 42.4 | |
| BBH | 46.8 | 56.3 | 44.1 | 65.2 | |
| TriviaQA‡ | 59.6 | 51.1 | 65.1 | 66.3 | |
| Code | HumanEval | 28.0 | 42.1 | 29.9 | 48.1 |
| MBPP | 48.7 | 57.1 | 43.2 | 63.8 | |
| Math | GSM8K | 34.0 | 79.1 | 41.1 | 77.4 |
| MATH | 8.5 | 42.6 | 17.1 | 45.3 | |
| CMath | - | 80.0 | 58.4 | 81.1 | |
| Chinese | C-Eval | - | 75.0 | 60.3 | 77.2 |
| CMMLU | - | 75.0 | 64.3 | 78.2 |
| Model | #Total Params | #Activated Params | Context Length | Download Link |
|---|---|---|---|---|
| Moonlight-16B-A3B | 16B | 3B | 8K | 🤗 Hugging Face |
| Moonlight-16B-A3B-Instruct | 16B | 3B | 8K | 🤗 Hugging Face |
transformers==5.13.0. Recommended environment: python=3.10+,
torch>=2.1 (verified on torch==2.6.0+cu124), transformers==5.13.0. Both snippets below
were re-run end-to-end against this fork and the verified outputs are shown after each.modeling_deepseek.py
patches described above if you want to run the base model on transformers==5.13.0:1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "melhoushi/Moonlight-16B-A3B-Instruct-transformers-5.13.0" # this fork; the base variant follows the same pattern
4model = AutoModelForCausalLM.from_pretrained(
5 model_name,
6 torch_dtype="auto",
7 device_map="auto",
8 trust_remote_code=True,
9)
10tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
11
12prompt = "1+1=2, 1+2="
13inputs = tokenizer(prompt, return_tensors="pt", padding=True, truncation=True).to(model.device)
14generated_ids = model.generate(**inputs, max_new_tokens=100)
15response = tokenizer.batch_decode(generated_ids)[0]
16print(response)1+1=2, 1+2=3, 2+1=3, 2+2=4, 3+1=4, 3+2=5, 3+3=6, 4+1=5, 4+2=6, 4+3=7, 4+4=8, 5+1=6, 5+2=7, 5+3=8, 5+4=9, 5+5=10,apply_chat_template(..., return_tensors="pt") returns a BatchEncoding in
transformers>=5, not a bare tensor, so we pass return_dict=True and unpack it into
generate(**inputs, ...):1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "melhoushi/Moonlight-16B-A3B-Instruct-transformers-5.13.0"
4model = AutoModelForCausalLM.from_pretrained(
5 model_name,
6 torch_dtype="auto",
7 device_map="auto",
8 trust_remote_code=True,
9)
10tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
11
12messages = [
13 {"role": "system", "content": "You are a helpful assistant provided by Moonshot-AI."},
14 {"role": "user", "content": "Is 123 a prime?"},
15]
16inputs = tokenizer.apply_chat_template(
17 messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
18).to(model.device)
19generated_ids = model.generate(**inputs, max_new_tokens=500)
20response = tokenizer.batch_decode(generated_ids)[0]
21print(response)<|im_system|>system<|im_middle|>You are a helpful assistant provided by Moonshot-AI.<|im_end|><|im_user|>user<|im_middle|>Is 123 a prime?<|im_end|><|im_assistant|>assistant<|im_middle|>To determine if 123 is a prime number, we need to check if it has any divisors other than 1 and itself. A prime number is a number greater than 1 that has no positive divisors other than 1 and itself.
Let's check for divisibility by smaller prime numbers:
1. **Divisibility by 2**: 123 is an odd number, so it is not divisible by 2.
2. **Divisibility by 3**: Sum the digits of 123: \(1 + 2 + 3 = 6\). Since 6 is divisible by 3, 123 is also divisible by 3.
Since 123 is divisible by 3, it is not a prime number.<|im_end|>@misc{liu2025muonscalablellmtraining,
title={Muon is Scalable for LLM Training},
author={Jingyuan Liu and Jianlin Su and Xingcheng Yao and Zhejun Jiang and Guokun Lai and Yulun Du and Yidao Qin and Weixin Xu and Enzhe Lu and Junjie Yan and Yanru Chen and Huabin Zheng and Yibo Liu and Shaowei Liu and Bohong Yin and Weiran He and Han Zhu and Yuzhi Wang and Jianzhou Wang and Mengnan Dong and Zheng Zhang and Yongsheng Kang and Hao Zhang and Xinran Xu and Yutao Zhang and Yuxin Wu and Xinyu Zhou and Zhilin Yang},
year={2025},
eprint={2502.16982},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.16982},
}