This dataset is tailored for fine-tuning embedding models in Retrieval-Augmented Generation (RAG) setups. It consists of 7,000 question-context pairs translated into Arabic, sourced from NVIDIA's 2023 SEC Filing Report.
The dataset is designed to improve the performance of embedding models by providing positive samples for financial question-answering tasks in Arabic.
This dataset is the Arabic version of the original… See the full description on the dataset page:
https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-finanical-rag-embedding-dataset.