Fork of bigcode/bigcodebench v0.1.4.
canonical_solution and test are ported to current numpy / pandas / scikit-learn
so that the reference solutions pass on a modern stack. Prompts
(complete_prompt, instruct_prompt, code_prompt) are unchanged, so the task a
model is asked to solve is identical to upstream.
Groundtruth pass rate: 1.000 over all 1140 tasks.