This dataset documents failure cases (“blind spots”) observed while testing the base language model :contentReference[oaicite:0]{index=0}.
The goal of this dataset is to identify scenarios where the model produces incorrect, misleading, or unreliable outputs when responding to prompts.
The dataset includes three main fields: