This is the dataset for DetailBench, which answers the question: "How good are current LLMs at finding small errors, when they are not explicitly asked to do so?"
article_title: Name of the Wikipedia article the data is from
original_text: Original excerpt from the given Wikipedia article
modified_text: Modified version of the original text with a single error (one changed number) introduced
original_number: The original number from the text… See the full description on the dataset page:
https://huggingface.co/datasets/xeophon/detailbench.