This dataset contains truncated tokenized protein sequences and their corresponding 3Di structure as stated in the Foldseek paper.
Redundancy reduction and data sequence filtering was performed by Dr. Michael Heinzinger and Prof. Dr. Martin Steinegger.
The tokenizer used to encode the sequences can be found here
More Information needed