This dataset aims to be a base template for fine-tuning embedding models for enhanced retrieval performance in RAG pipelines.
It has been generated locally using Mistral:7B on Ollama using a simple prompt that prompts the model to generate 5 questions for each document chunk of Apple's Environmental Progress Report 2024