Private dataset backing the PRISM2 family of antibody language models trained
at the Romero Lab (Duke University). Bundles the OAS immune-repertoire and
SAbDab structural corpora used for the v6 / v7 unified VDJ runs, plus the
V/D/J gene vocabulary files.
All sequence data has been pre-filtered for the production germline-distance
threshold used in PRISM2 training, so what you load here is what the model
trained on — no extra filtering required.… See the full description on the dataset page:
https://huggingface.co/datasets/RomeroLab-Duke/prism2-antibody-dataset.