Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen2.5-72B-Instruct-Malaysian – AI Model by Scicom-intl | AlphaNeural AI
You can deploy this model and start earning money today!
Scicom-intl
/
Qwen2.5-72B-Instruct-Malaysian
like
0
safetensors
qwen2
Scicom-intl/Malaysian-Instructions
Qwen/Qwen2.5-72B-Instruct
finetune
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen2.5-72B-Instruct-Malaysian
SFT LoRA
Qwen/Qwen2.5-72B-Instruct
on
Scicom-intl/Malaysian-Instructions/commit/288b358a57765a735d588f73e5e6c212c81429bd
Dense LoRA SFT done using DeepSpeed Zero3 HF Trainer.
Multipacking variable length 16384 context length, with global batch size of 32, so global total tokens is 524288.
All linear layers with rank 256 with alpha multiply by 2.0
Liger fused cross entropy.
1e-4 learning rate, 50 warmup, 3 epoch only.
We only upload the best model
Source code
Source code at
https://github.com/Scicom-AI-Enterprise-Organization/small-ablation/blob/main/malaysian-sft
Acknowledgement
Special thanks to
https://www.scitix.ai/
for H100 Node!