Views
No views yet
Qwen/Qwen3-4B-Base continued-pretrained on the train split of
JackHsieh/statML-arxiv-40M-20M
(L=4096) — the thoughtless baseline (plain next-token prediction, no thoughts), best cell
of the prestar 4B-base-thoughtless lr × replay sweep.model = AutoModelForCausalLM.from_pretrained("JackHsieh/Qwen3-4B-Base-thoughtless.statml-arxiv-40M-20M.best")