MinorBench: A Benchmark for Child-Safety in LLMs
Dataset Summary
MinorBench is a benchmark designed to evaluate whether large language models (LLMs) respond to questions that may be inappropriate for children, particularly in an educational setting. It consists of 299 prompts spanning various sensitive topics, assessing whether models can appropriately filter or refuse responses based on child-friendly assistant roles.
The benchmark pairs each prompt with one of four… See the full description on the dataset page: https://huggingface.co/datasets/govtech/MinorBench.