Generated by Faker, contains data for testing searching model numbers using vectorized model numbers.
This dataset does NOT include embeddings. See other datasets for a smaller sample, and one with embeddings.
Size: 50,000 entries.
Columns:
brand
model_number
model_name
year
randomdata: between 1000 and 2000. Append to model_number if the faked value is under 6 characters.
model_search: remove some characters (see below) from model_number. This used for creating… See the full description on the dataset page:
https://huggingface.co/datasets/blade57/ModelNumbers4Searching_Full.