Some bad data discovered in the popular GSM8K and SVAMP LLM benchmarking datasets.
These examples have incorrect answers in the corresponding math problem benchmark dataset, and should not be used to evaluate AI models.
We detected this bad data automatically using Cleanlab's Trustworthy Language Model. TLM's estimated trustworthiness score for each example is also provided.
Question: After scoring 14 points, Erin now has three times… See the full description on the dataset page:
https://huggingface.co/datasets/Cleanlab/bad_data_gsm8k_svamp.csv.