The definitive retrieval corpus for PEX10 / peroxisomal biogenesis disorder research.
A pre-built ChromaDB vector database containing:
98,734 indexed text chunks from 835 curated PubMed Central (PMC) biomedical papers + 1,495 structured curated entries (1,016 ClinVar variants + 198 truncation consequence cards + 281 ESMFold variant structural analyses),
embedded with NVIDIA's state-of-the-art Llama-Nemotron-Embed-1B-v2⦠See the full description on the dataset page:
https://huggingface.co/datasets/SkyWhal3/PEX10-RAG-Nemotron.