This dataset contains 16,181 prompt-response evaluation pairs from OpenAI's gpt-oss-20b model, generated as part of a large-scale red-teaming effort for the Kaggle Red-Teaming Challenge.
The evaluations are sourced from 15+ distinct public red-teaming and safety datasets. Each record includes the original prompt, the model's response, token counts, a harm category classification, the source dataset, and binary flags for… See the full description on the dataset page: https://huggingface.co/datasets/ChestnutKurisu/gpt-oss-20b-red-teaming-evals.