This dataset contains 10 diverse failure cases (“blind spots”) observed when probing the base language model LiquidAI/LFM2.5-1.2B-Base.
input: the prompt
expected_output: the expected correct/ideal response
model_output: the model’s observed output
error_type: category of failure… See the full description on the dataset page:
https://huggingface.co/datasets/makda-tsegazeab/lfm2-1.2b-blindspots.