The full experimental workflow is documented in the Jupyter notebook: Base_Model_Evaluation.ipynb
The notebook contains the complete pipeline used to load the model, generate outputs, evaluate behavioral patterns, and construct a dataset capturing the model’s blind spots.
This project evaluates the behavior of a base large language model when performing different types of tasks such as: