[Phi-4-Super-o1 finetuned] from Microsoft's Phi-4 is a state-of-the-art open model developed with a focus on responsible problem solving and advanced reasoning capabilities. Built upon a diverse blend of synthetic datasets, carefully filtered public domain websites, and high-quality academic books and Q&A datasets, Phi-4-Super-o1 ensures that small, capable models are trained with datasets of exceptional depth and precision.
Phi-4-Super-o1 adopts a robust safety post-training approach using open-source and in-house synthetic datasets. This involves a combination of SFT (Supervised Fine-Tuning) and iterative DPO (Direct Preference Optimization) techniques, ensuring helpful and harmless outputs across various safety categories.
Let me know if you’d like further edits or refinements!
This is a merge of pre-trained language models created using
mergekit.
This model was merged using the
Model Stock merge method using
unsloth/phi-4 as a base.
1models:
2 - model: unsloth/phi-4
3 - model: prithivMLmods/Phi-4-o1
4 - model: prithivMLmods/Phi-4-Math-IO
5 - model: prithivMLmods/Phi-4-QwQ
6 - model: Pinkstack/SuperThoughts-CoT-14B-16k-o1-QwQ
7 - model: mudler/LocalAI-functioncall-phi-4-v0.3
8 - model: bunnycore/Phi-4-RP-V0.2
9 - model: prithivMLmods/Phi-4-Empathetic
10merge_method: model_stock
11base_model: unsloth/phi-4
12parameters:
13 normalize: false
14 int8_mask: true
15dtype: bfloat16
16tokenizer_source: "unsloth/phi-4"
17