Views
No views yet
prot, DNA, RNA, other). Files are chunked/merged from 1k-folders and cleaned for consistent CSV schema.
Current data based on only 120K proteins.3di_chains_chaintag_filtered_prot_only.csv filtered out repeated protein and DNA RNA chain. Main file used for SFT_all contains all chains (including D-amino, RNA, DNA, and any others)._prot_only contains protein chains only (L- and D-amino acids treated as protein)._sample_1000 is a 1,000-row random sample for quick inspection.index — global row index from the source CSV (0-based)pdb_id — 4-character PDB code (e.g., 9B4J)chain_id — chain identifier (alphanumeric, may include digits/letters)aa_seq — amino-acid sequence (when available)threeDi_seq — Foldseek 3Di token sequence (Used for sft)combined_seq — helper concatenation (3Di/AA) used upstreamseq_len — chain sequence length (prefer AA length; else derived)chunk — source folder name (e.g., 1000_1999)path — absolute path to the structure file usedpolymer_class — one of prot, DNA, RNA, otherprot)_all)