This repository is responsible for the entire process of collecting, cleaning, and standardizing Orthohantavirus genomic data for the HantaBERT project. The pipeline automates data extraction from NCBI GenBank to produce a ready-to-use dataset for machine learning.
Extraction Automation: Uses Biopython to fetch thousands of RNA sequences (S, M, L) and related metadata in batches from the NCBI database.
Multi-task Labeling:… See the full description on the dataset page:
https://huggingface.co/datasets/HantaBERT/Orthohantavirus-Genome-Atlas.