The dataset consist of questions in the 3 styles "implicit", "search string" and "standard questions".
As well as of the additional Positive and Hard Negative document example. Additionally all intermediate data was kept.
Benchmarks of RAG Pipelines showed differing results for the different question styles.
The Generation-Pipelone is based on vLLM and Mixtral in 4bit quant. All scripts:
https://github.com/SebastianBodza/Embedding_Training
One… See the full description on the dataset page:
https://huggingface.co/datasets/SebastianBodza/synthetic_RAG_dataset_ger_de_v02.