This is a binary classification version of the finetuning/evaluation datasets introduced in
the freeform text annotations of the greenbeing-proteins dataset.
Proteins from UniProtKB (knowledge base), from select food crops and related species.
Amino acid sequences use IUPAC-IUB codes where letters A-Z map to amino acids.
The "query" field currently contains 18 possible labels from the annotation, and the "label" field is the binary class (0 =… See the full description on the dataset page:
https://huggingface.co/datasets/monsoon-nlp/greenbeing-binary.