Views
No views yet
Qwen/Qwen2.5-1.5B-Instruct, Sift-1B strips away all conversational fluff ("Sure! Here is your JSON:") and outputs strict, machine-readable JSON on the very first attempt.Qwen/Qwen2.5-1.5B-Instructq4_k_m)SanatanSinghVishen/sift-1b-ggufSanatanSinghVishen/sift-1b-dpoSanatanSinghVishen/sift-1b-sft| Evaluation Metric | Base Model (Qwen2.5-1.5B) | Sift-1B (SFT) | 🏆 Sift-1B (DPO Golden) | Delta vs Base |
|---|---|---|---|---|
| Tool Selection Accuracy | 70.0% | 98.0% | 100.0% ✅ | +30.0% |
| Parameter Extraction Accuracy | 34.0% | 80.0% | 88.0% ✅ | +54.0% |
| JSON Parse / Validity Rate | 96.0% | 98.0% | 100.0% ✅ | +4.0% |
| Zero Markdown / Fluff Rate | 76.0% | 100.0% | 100.0% ✅ | +24.0% |
| Zero Hallucination Rate | 100.0% | 100.0% | 100.0% ✅ | 0% Hallucinations |
| Average Latency (TTFT) | 2,277 ms | 1,734 ms | 1,714 ms ⚡ | 25% Faster |
q4_k_m (4-bit medium K-quantization)940.4 MB (0.94 GB)n_ctx): 32,768 tokensn_embd): 1,536n_ff): 8,960n_head): 12n_head_kv): 2 (Grouped-Query Attention / GQA)1e-6rope_theta): 1,000,000.0151,936 tokens (ChatML format)load_in_4bit=True)q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj2.0e-4 (Cosine schedule, warmup_ratio=0.05)adamw_8bit0.1 | Loss Type: sigmoid5.0e-6 (Cosine schedule, warmup_ratio=0.1)1# Run directly from Hugging Face Hub:
2ollama run hf.co/SanatanSinghVishen/sift-1b-gguf