A mistral-7B-v0.3-instruct is trained using 17K harmless and helpful dataset (see harmless-base, helpful-base) using Parameter Efficient Fine-tuning with LoRA rank of 16.
Trained on a single H100 SXM for ~3.5-4 hours.
See sample generations from finetuned model: saferlhf_prompts_10_sample_generations.json
Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More Information Needed]
Downstream Use [optional]
[More Information Needed]
Out-of-Scope Use
[More Information Needed]
Bias, Risks, and Limitations
[More Information Needed]
Recommendations
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.