This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the heavyweight StarCoder 15B base foundational model.
This specific partition documents the absolute scaling limits of unaligned foundational weights inside conversational benchmarking loops, establishing a massive baseline for… See the full description on the dataset page:
https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder_15b.