This dataset is a part of an efficiency evaluation framework for LLM-generated code.
It is based on Mercury, a competitive programming benchmark, with the following improvements:
test generation facilities and extended test suites
corrected test generation, input conversion and solution evaluation functions
enhanced sandbox with additional library imports and adjusted time/recursion limits
See this repository on Github for more details.