Local Code Arena Telemetry: MBPP Benchmark on DeepSeek R1 1.5B
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the DeepSeek R1 1.5B distilled reasoning architecture.
This specific run establishes the performance boundaries of lightweight reasoning models under strict execution time limits on consumer hardware.
📊 Core Performance… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-deepseek-r1_1.5b.