This dataset is used to train a Seq2Seq model designed to fix ingredient lists of Open Food Facts products.
Products were extracted from the Open Food Facts database (JSONL) along the lang and the list of ingredients. These products were selected
in respect of some criteria:
20 to 40% unknown ingredients computed during the Ingredient Extraction Analysis,
No duplicate in the list of ingredients,
No duplicate with the spellcheck-benchmark