This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the flagship StarCoder2 15B base foundational model.
This specific partition documents the final limits of scaling raw, unaligned foundational weights inside conversational evaluation loops, establishing an absolute baseline for… See the full description on the dataset page:
https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder2_15b.