Welcome to the curated DNA sequence dataset, automatically gathered from NCBI using the Enigma2 pipeline. This repository provides ready-to-use CSV and Parquet files for downstream machine-learning and bioinformatics tasks.
A collection of topic-specific DNA sequence sets (e.g., BRCA1, TP53, CFTR) sourced directly from NCBI’s Nucleotide database.
Predefined Entrez queries (gene… See the full description on the dataset page:
https://huggingface.co/datasets/shivendrra/EnigmaDataset.