This dataset contains prompts and responses from various models, including accepted and rejected responses based on specific criteria. The dataset is designed to help in the study and development of adversarial attacks and alignment in reinforcement learning from human feedback (RLHF).
Prompts: Various prompts used to elicit responses from models.
Accepted Responses: Responses that were accepted based on specific… See the full description on the dataset page:
https://huggingface.co/datasets/yaswanth-iitkgp/Adversarial_Attacks_Alignment_Dataset.