GRPO-finetuned LoRA adapters on Qwen3-14B for the NYCU DL HW3 Chinese-MCQ task.
Warm-started from HW1 SFT adapter (Claire0730/DL_HW1_SFT, qwen3-14b-qlora-s42),
trained with TRL GRPO (4-bit QLoRA) on mined hard/uncertain questions.
grpo-14b-hard-s42/ — GRPO v1 (549 hard Qs, 140 steps, lr 1e-6)
grpo-14b-v2/ — GRPO v2 (~1200 hard Qs, 250 steps, lr 2e-6, completion 768)
data/ — GRPO training data (hard sets + full pool)