Views
No views yet
zhengbang0707/Qwen3.5-9B-posthoc on GSM8K. Targets contain the final answer only (no chain-of-thought). Training uses Qwen3.5's multimodal conditional-generation architecture, the native chat template with enable_thinking=False, BF16 parameters, and 4-GPU FSDP. The final model was consolidated without intermediate optimizer checkpoints.f5abed0c7a87aeaa678d64db64a8816b16c90f34openai/gsm8k, main train split; validation loss on the held-out splitus.anthropic.claude-sonnet-4-6 in us-east-2AutoModelForMultimodalLM and AutoProcessor from Transformers. Keep the Qwen3.5 native processor/chat-template path and set enable_thinking=False for no-CoT inference.