ScentSet is a synthetic dataset containing 572,293 entries and approximately 15 million tokens. Each entry is a short natural language description in simple english of a smell, often followed by a hint or guess about its source. The dataset is designed to support machine learning research in scent recognition, classification, and multimodal representation learning.
{"text": "There's a bright… See the full description on the dataset page:
https://huggingface.co/datasets/sixf0ur/ScentSet.