This dataset contains 10 evaluated test cases documenting the blind spots and failure modes
of Qwen/Qwen2.5-1.5B when applied to healthcare and medical AI prompts. 9 out of 10 prompts
revealed distinct failure modes — ranging from dangerous clinical misinformation to broken
multilingual output. One prompt… See the full description on the dataset page: https://huggingface.co/datasets/Nirjhor/llm-healthcare-blindspots.