This model is a fine-tuned version of
meta-llama/Meta-Llama-3-8B on UltraChat SFT with the respond token, CoCoNoT refusals with the refuse token, and CoCoNoT's contrast data as SFT data with the multiple refusal tokens for each of the five categories. Note that this model is not the model found in the paper the original models are not able to be released due to corporate legalities.
For generating a output from this model, please refer to the code found in
repo in the
coconot_eval folder.