Described real images using BLIP (Bootstrapping Language-Image Pre-training)
Generated Stable Diffusion images using BLIP descriptions
Found similar Midjourney images based on BLIP descriptions
This approach ensured real and AI-generated images were as similar as possible, differing only in their origin.
The three models were then distilled into a small ViT model with 11.8 Million Parameters, combining their learned features for more efficient detection.