Here we present the variant level embeddings for large-scale genetic analyis as described in 'Incorporating LLM Embeddings for Variation Across the Human Genome,' based on curated annotations using high quality functional data from FAVOR, ClinVar, and GWAS Catalog. We currently present embeddings using either OpenAI's text-embedding-3-large (3072-dimensional) or Qwen's Qwen3-Embedding-0.6B (1024-dimensional) models.
Genetic variants are identified with… See the full description on the dataset page:
https://huggingface.co/datasets/LiLabUNC/Variant-Foundation-Embeddings.