Views
No views yet
[!IMPORTANT] These Llama models are initialized from Qwen models with the embedding layer adapted to fit the Llama tokenizer. This enables efficient cross-model family knowledge transfer.
| Model | Base Model | Base + MemDec |
|---|---|---|
| Llama3-8B | 5.96 | 4.46 |
| Llama3-70B | 4.90 | 4.07 |
| Model | Base Model | Base + MemDec |
|---|---|---|
| Llama3.1-8B | 5.88 | 4.42 |
| Llama3.1-70B | 4.89 | 4.06 |
| Model | Base Model | Base + MemDec |
|---|---|---|
| Llama3.2-1B | 8.23 | 5.11 |
| Llama3.2-3B | 6.83 | 4.76 |
1@article{cao2025memory,
2 title={Memory decoder: A pretrained, plug-and-play memory for large language models},
3 author={Cao, Jiaqi and Wang, Jiarui and Wei, Rubin and Guo, Qipeng and Chen, Kai and Zhou, Bowen and Lin, Zhouhan},
4 journal={arXiv preprint arXiv:2508.09874},
5 year={2025}
6}