Youtu-LLM-2B-Base Blind Spots Dataset
What is this?
I tested a small AI language model called Youtu-LLM-2B-Base (made by Tencent) to find places where it gives wrong or strange answers. I gave it 50 different questions and kept the 25 cases where it clearly failed.
This dataset contains those 25 failures — the question I asked, what the correct answer should be, what the model actually said, and why it was wrong.