This model is a scratch-built, highly optimized implementation of the CLIP architecture, developed as part of an Academic Research Project.
It achieves 2.46x faster inference speed (Latency: 21ms vs 52ms) compared to the standard OpenAI CLIP model on consumer hardware (RTX 3050 Ti), while maintaining 97.7% Zero-Shot Accuracy.