ProCyon-Split is a multimodal foundation model for protein phenotypes, which combines a large language model with protein encoders to support inputs of interleaved free text and proteins.
In contrast to ProCyon-Full, this model is instruction-tuned using the training split of the
ProCyon-Instruct dataset to
enable rigorous model evaluation on held-out protein-phenotype pairs.
For more information on the model design, training, and validation, please see the
overview page or the
paper.
Additional versions of the model are available as
ProCyon-Full and
ProCyon-Bind.