Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step3000 – AI Model by huggingFacing | AlphaNeural AI
You can deploy this model and start earning money today!
huggingFacing
/
qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step3000
like
0
transformers
safetensors
qwen2
text-generation
qwen2.5
knowledge-distillation
gkd
liftkd
conversational
en
zh
Qwen/Qwen2.5-1.5B-Instruct
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen2.5 7B to 1.5B LiftKD V8 Bilingual 100K - Epoch 3
This is the cumulative epoch-3 checkpoint of a Qwen2.5-1.5B-Instruct student distilled from Qwen2.5-7B-Instruct.
Method: LiftKD V8, fully on-policy GKD JSD with normalized gap gate
Data: 100K bilingual English/Chinese instruction mixture, including 18.75% mathematics
Sequence limits: 384 prompt tokens, 512 total tokens, 128 generated tokens
Precision: BF16 full-parameter training with DeepSpeed ZeRO-2
Global batch size: 64
Seed: 10
The training mixture was internally deduplicated. Evaluation-set decontamination was not performed.