This dataset contains 12 diverse stress-test data points identifying the logical, physical, and ethical "blind spots" of the Gemma-2-2b base model. By using leading completions, we uncover how the model's raw weights handle reasoning, bias, and instruction drift without the safety layers of instruction tuning.
Model Name: google/gemma-2-2b
Parameters: 2.5 Billion
Type: Base Model (Causal Language Model)… See the full description on the dataset page:
https://huggingface.co/datasets/Ahmed-Nasri/Fatima-Fellowship-challenge.