Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
DL_hw3 – AI Model by Edward1239 | AlphaNeural AI
You can deploy this model and start earning money today!
Edward1239
/
DL_hw3
like
0
peft
safetensors
qwen
grpo
lora
multiple-choice
text-generation
conversational
Qwen/Qwen2.5-14B-Instruct
adapter
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
DL HW3 GRPO LoRA Adapter
This repository contains the final LoRA adapter for DL HW3: Reasoning LLM Step 3 with GRPO.
Base Model
Qwen/Qwen2.5-14B-Instruct
Method
The model was initialized from my Step2 SFT LoRA adapter and further optimized using GRPO.
Final Adapter
outputs/grpo_hw2best_balanced_30steps_lr2e8
Training Data
The final GRPO run used a balanced version of HW2_.csv:
A: 261
B: 261
C: 261
D: 261
Inference
Final inference uses score-only A/B/C/D log-softmax scoring.
Public Leaderboard
Public LB score: around 0.71