A lightly patched fork of FudanSELab/ClassEval,
the 100-task class-level Python code generation benchmark, fixed so that it
still runs correctly on a current Python and a current NumPy.
The benchmark itself is unchanged. Every patch either repairs a test that no
implementation could pass, or repairs the reference solution. No task was made
easier, no prompt (skeleton) was touched, and nothing about what a model is
asked to write has changed.… See the full description on the dataset page:
https://huggingface.co/datasets/ilintar/ClassEval.