Adversarial Evaluation of Qwen1.5-1.8B-Chat-GGUF
Model Overview
This dataset contains failure points identified during an adversarial evaluation of the Qwen/Qwen1.5-1.8B-Chat-GGUF model.
Logical Reasoning: The model struggles with syllogistic reasoning (e.g., 'all bloops are blips') and classic mathematical word problems, often… See the full description on the dataset page:
https://huggingface.co/datasets/Mohamed682004/qwen1.5-adversarial-eval.