This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the ultra-lightweight DeepSeek Coder 1.3B model.
This specific run establishes the absolute maximum throughput envelope of our local hardware setup while tracking the accuracy trade-offs of legacy, lightweight code specialists.… See the full description on the dataset page:
https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-deepseek-coder_1.3b.