SmellBench: Towards Fine-Grained Evaluation of Code Agents on Refactoring Tasks
Dataset Summary
SmellBench is a benchmark designed to evaluate whether code agents can detect and refactor bad code (code smells). Each instance represents a validated code smell injection case constructed from real-world open-source repositories, enabling fine-grained assessment of code agents' refactoring capabilities.