A multi-source protein complex dataset for self-supervised preference training.
Every cross-species swap is ortholog-aligned (swap human_A with yeast_A —
the orthologous gene's product — not a random yeast protein) and labelled
with three sequence-identity metrics + TimeTree species divergence time.
Identity filter (v3.1): L2/L3/L4 are restricted to ortholog pairs whose
BLAST-style identity ≥ 0.80 on both chain A and chain B. Below this the
"swap" is too divergent to… See the full description on the dataset page:
https://huggingface.co/datasets/wjiaqi/evo-final.