The dataset is a synthetically generated collection of 10,000 online courses created specifically for content-based recommendation tasks.
It was generated using the mistralai/Mistral-7B-Instruct-v0.2 language model, ensuring consistent structure, high semantic quality, and broad topical coverage.
The dataset is well suited for experiments involving embeddings, semantic similarity, and text-based recommendation systems.
During… See the full description on the dataset page:
https://huggingface.co/datasets/Daniel-1109/CoursePilot.