Local Code Arena Telemetry: MBPP Benchmark on Qwen 2.5 Coder 7B
This dataset contains raw, end-to-end evaluation metrics, generation outputs, execution logs, and abstract syntax tree (AST) structural performance values collected from executing the Mostly Basic Python Problems (MBPP) benchmark.
The evaluation was performed completely offline using a local consumer high-performance compute profile to track execution fidelity without cloud bias or network variance.
📊 Core… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-qwen-7b.