This dataset was created to evaluate the model's behavior on reasoning traps and ambiguous instructions. The goal is to identify systematic weaknesses (blind spots) that can be addressed through targeted fine-tuning.
The model tested in this dataset is tiny-aya-base, developed by Cohere.
The model is designed as a lightweight multilingual instruction-following language model.
Model link:
https://huggingface.co/CohereLabs/tiny-aya-base
This model was evaluated on a small… See the full description on the dataset page:
https://huggingface.co/datasets/bit-wander/tiny-aya-base-blindspot.