A curated, multi-language benchmark for code fault localization — the task
of identifying which source-code symbols (functions, methods, classes, etc.)
need to be modified to resolve a given issue.
Built on top of SWE-bench Verified and
SWE-bench Multilingual,
this dataset augments each instance with ground-truth locator annotations
extracted by tree-sitter–based code chunking and patch analysis.
Property
Value… See the full description on the dataset page:
https://huggingface.co/datasets/fishmingyu/codenib-base-dataset.