JB-Completions Dataset: Base Model Safety Evals
Overview
JB-Completions is a dataset designed for evaluating the harmfulness of base language models (i.e., completion/non-instruction-fine-tuned LLMs). This dataset contains pairs of harmful prompts and their corresponding completions, allowing researchers to assess how base models respond to potentially harmful inputs. See our paper on Safety Pretraining for more details!