Views
No views yet
| Developers | Microsoft Research |
| Description | phi-4 is a state-of-the-art open model built upon a blend of synthetic datasets, data from filtered public domain websites, and acquired academic books and Q&A datasets. The goal of this approach was to ensure that small capable models were trained with data focused on high quality and advanced reasoning.phi-4 underwent a rigorous enhancement and alignment process, incorporating both supervised fine-tuning and direct preference optimization to ensure precise instruction adherence and robust safety measures |
| Architecture | 14B parameters, dense decoder-only Transformer model |
| Inputs | Text, best suited for prompts in the chat format |
| Context length | 16K tokens |
| GPUs | 1920 H100-80G |
| Training time | 21 days |
| Training data | 9.8T tokens |
| Outputs | Generated text in response to input |
| Dates | October 2024 – November 2024 |
| Status | Static model trained on an offline dataset with cutoff dates of June 2024 and earlier for publicly available data |
| Release date | December 12, 2024 |
| License | MIT |
phi-4 is best suited for prompts using the chat format as follows:1<|im_start|>system<|im_sep|>
2You are a medieval knight and must provide explanations to modern people.<|im_end|>
3<|im_start|>user<|im_sep|>
4How should I explain the Internet?<|im_end|>
5<|im_start|>assistant<|im_sep|>transformers1import transformers
2
3pipeline = transformers.pipeline(
4 "text-generation",
5 model="microsoft/phi-4",
6 model_kwargs={"torch_dtype": "auto"},
7 device_map="auto",
8)
9
10messages = [
11 {"role": "system", "content": "You are a medieval knight and must provide explanations to modern people."},
12 {"role": "user", "content": "How should I explain the Internet?"},
13]
14
15outputs = pipeline(messages, max_new_tokens=128)
16print(outputs[0]["generated_text"][-1])