This dataset contains 11 prompts where the base model
unsloth/Qwen3-0.6B-Base
produces incorrect or undesirable outputs.
Each row has:
tag: rough failure category (math, factual, reasoning, instruction, memory, safety, lexical)
input: the prompt given to the model
expected: the answer a knowledgeable human would give
model_output: the raw model output from the model