A dataset of 4,073 operons across 11 bacterial genomes species.
The operon annotations have been extracted from Operon DB and the genome DNA sequences have been extracted from GenBank. Each row contains whole bacterial genome, represented
by a list of DNA sequences from different contigs.
We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous operons and only kept genomes with at least 9 known… See the full description on the dataset page:
https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-dna.