Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-8b-dflash-nemotron-codealpaca-greedy-regen-distill – AI Model by JensenYuan | AlphaNeural AI
You can deploy this model and start earning money today!
JensenYuan
/
qwen3-8b-dflash-nemotron-codealpaca-greedy-regen-distill
like
0
safetensors
qwen3
custom_code
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen3-8B DFlash Draft Model (Distillation)
Base model
: Qwen/Qwen3-8B
Training method
: DFlash with teacher-student distillation loss
Checkpoint
: epoch_1_step_50000
Dataset
: nemotron + codealpaca (greedy regen)
Hyperparameters
:
batch_size: 2
learning_rate: 3e-4
loss_type: distill
loss_decay_gamma: 7.0
block_size: 16
num_epochs: 2