The benchmark (v5) is composed of 152 lists of ingredients extracted from the Open Food Facts database, then corrected to ensure that each ingredient is recognized.
Its purpose is to evaluate the Spellcheck on correcting products list of ingredients in respect of the OFF guidelines.
A portion of the data was synthetically generated using OpenAI-GPT3.5-Turbo prompted for this task.
Once composed, the benchmark was then checked with the annotation tool Argilla… See the full description on the dataset page:
https://huggingface.co/datasets/openfoodfacts/spellcheck-benchmark.