Dataset made of diverse genomes available on NCBI and coming from ~850 different species.
Test and validation are made of 50 species each. The rest of the genomes are used for training.
Default configuration "6kbp" yields chunks of 6.2kbp (100bp overlap on each side). Similarly,
the "12kbp"configuration yields chunks of 12.2kbp. The chunks of DNA are cleaned and processed so that
they can only contain the letters A, T, C, G and N.