This model is a fine-tuned version of Qwen/Qwen3-4B-Instruct-2507 using Direct Preference Optimization (DPO) via the Unsloth library.
This repository contains the full-merged 16-bit weights. No adapter loading is required.
Training Objective
This model has been optimized using DPO to align its responses with preferred outputs, focusing on improving reasoning (Chain-of-Thought) and structured response quality based on the provided preference dataset.
Training Configuration
Base model: Qwen/Qwen3-4B-Instruct-2507
Method: DPO (Direct Preference Optimization)
Epochs: 1
Learning rate: 1e-07
Beta: 0.1
Max sequence length: 1024
LoRA Config: r=8, alpha=16 (merged into base)
Usage
Since this is a merged model, you can use it directly with transformers.