You probably do not need this unless you are training your own IP Adapters.
Modified version of the vision encoder of
CLIP-ViT-H-14-laion2B-s32B-b79K to handle 448 x 448 inputs
vs the original 224 x 224 inputs. It will probbaly not work for classification (as is), but will DIP work for for IP+ adapters that use CLIP-ViT-H, though they will need to be
fine tuned a little more.