Dataset designed for fine-tuning a small LLM (e.g. gemma-3-270m) to extract structured data from text in a way which replicates a much larger LLM (e.g. gpt-oss-120b).
Purpose it to enable a fine-tuned small LLM to filter a large text dataset for food and drink-like items.
For example, take DataComp1B dataset and use the fine-tuned LLM to filter for food and drink related items.
{'sequence': 'A mouth-watering photograph captures a delectable… See the full description on the dataset page:
https://huggingface.co/datasets/bagherihamid/FoodExtract-1k.