Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen3-4B-Dense-Merged-Soup-30k – AI Model by ShourenWSR | AlphaNeural AI
You can deploy this model and start earning money today!
ShourenWSR
/
Qwen3-4B-Dense-Merged-Soup-30k
like
0
transformers
safetensors
qwen3
text-generation
model-soup
merge
conversational
ShourenWSR/Qwen3-4B-Dense-NoThink-30k
merge
ShourenWSR/Qwen3-4B-Dense-Think-30k
merge
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen3-4B-Dense-Merged-Soup-30k
A model soup (0.5/0.5 arithmetic weight average, no training) of two dense Qwen3-4B SFT checkpoints trained on the same 30k data, differing only in thinking mode:
ShourenWSR/Qwen3-4B-Dense-NoThink-30k (no_think-only SFT)
ShourenWSR/Qwen3-4B-Dense-Think-30k (think-only SFT)
Purpose
Single-model naive baseline ("merged-soup") in the PL-MoE rebuttal 4-way comparison: no_think-only / think-only / merged-soup / PL-MoE.
Merge details
Per-tensor arithmetic mean, weights 0.5 / 0.5
Accumulation in float32, cast back to original bf16
Config + tokenizer copied verbatim from the no_think parent (both parents byte-identical)
399 tensors averaged; key sets / shapes / dtypes verified identical before merge