Our ColonGPT is a standard multimodal language model, which contains four basic components: a language tokenizer, an visual encoder (🤗
SigLIP-SO), a multimodal connector, and a language model (🤗
Phi1.5). In this huggingface page, we provide a quick start for convenient of new users. For further details about ColonGPT, we highly recommend visiting our
homepage. There, you'll find comprehensive usage instructions for our model and the latest advancements in intelligent colonoscopy technology.