Views
No views yet
meta-llama/Llama-3.2-3B-Instructbeta=0)1e-06False / Falsesiyanzhao/Openthoughts_math_30k_opsdmax_steps=100, max_completion_length=1024per_device_train_batch_size=1, gradient_accumulation_steps=2, effective batch 8colocate, GPU memory utilization 0.35