Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
llm-sft-dpo-lora-repo20260215 – AI Model by centmount | AlphaNeural AI
You can deploy this model and start earning money today!
centmount
/
llm-sft-dpo-lora-repo20260215
like
0
safetensors
unsloth
qlora
sft
dpo
structured-output
Qwen/Qwen3-4B-Instruct-2507
finetune
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
llm-sft-dpo-lora-repo20260215 (SFT + DPO)
This model is fine-tuned for
Structured Output (JSON/XML/TOML)
using
SFT + DPO
.
Best Checkpoint: Step 120 (Validation Loss ~0.06, Accuracy ~0.98)
Methods
SFT
:
Cleaned dataset (removed markdown blocks, removed
<think>
tags).
Upsampled weak formats (TOML, XML).
Normalized TOML syntax.
DPO
:
Rejected samples included
<think>
tags, Markdown blocks, inline TOML, and CSV trailing commas.
Optimized with Beta=0.3, LR=1e-6.
Usage
Use with Unsloth or PEFT.