Preprocessed training assets for
PockLigGPT.
crossdocked/
├── crossdocked_clean_pocket_selfies_with_tokens.parquet
├── per_residue_index.parquet
└── per_residue_pack.npz.part-*
The CrossDocked training Parquet contains pocket metadata, SELFIES,
fixed-length token_ids, and offsets into the residue embedding stack.
The split NPZ contains… See the full description on the dataset page:
https://huggingface.co/datasets/pablovp8/PockLigGPT-training-data.