Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Nemotron-3-Nano-30B-A3B-bodhi-distil – AI Model by Krish2002 | AlphaNeural AI
You can deploy this model and start earning money today!
Krish2002
/
Nemotron-3-Nano-30B-A3B-bodhi-distil
like
0
safetensors
nvidia
nemotron
lora-merged
distillation
math
text-generation
conversational
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16
finetune
other
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Nemotron-3-Nano-30B-A3B BODHI distillation
LoRA fine-tune of
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16
on the
espressovi/BODHI-distillation
dataset, merged back into the base for standalone use.
Training
Base: NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16 (hybrid Mamba2 + attention + MoE)
Method: LoRA, r=64, alpha=128, attention modules only (q/k/v/o_proj)
Sequence length: 8192
Compute: 8x A100 80 GB, ZeRO-2 + grad checkpointing, 1 epoch (~860 steps)
License
Inherits the NVIDIA Open Model License from the base model.