This repository contains the ProCyon-Instruct used to train the ProCyon family of models.
Please see installation instructions on our GitHub repo for details on how to configure the dataset
for use with pre-trained ProCyon models. For additional technical details, please refer to our overview page or the paper.
The repository contains three top-level directories:
integrated_data/v1 - The primary component of the dataset: the amino acid sequences and associated phenotypes used for constructing… See the full description on the dataset page:
https://huggingface.co/datasets/mims-harvard/ProCyon-Instruct.