CLIP (Contrastive Language-Image Pre-training) performs zero-shot image classification by comparing image embeddings against text prompt embeddings. This repository packages CLIP evaluation on CIFAR-10 using PyTorch (CPU or GPU via OpenAI CLIP) and vLLM (GPU server with pooling runner, CPU HTTP client), exported and validated for AMD ROCm so it runs efficiently on AMD GPUs and CPUs.
This is based on the implementation of CLIP found here.
This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.
Task Overview
Task: Zero-shot image classification
Dataset: CIFAR-10 test split (10,000 images, 10 classes)
Output metrics: Zero-shot accuracy (%)
Quick Start: Before running this example, complete the main repository setup — see [Prerequisites](#prerequisites), [Configure System Paths](#2-configure-system-paths), and [Virtual Environments](#3-virtual-environments) in the main README.
Model variants: Default is base32 (openai/clip-vit-base-patch32 / ViT-B/32). Override with MODEL_VARIANT=base16|large14|large14-336.
vLLM note: vLLM's CLIP backend embeds one modality per request — text prompts and images are sent in separate API calls, then cosine similarity is computed client-side (same approach as the original evaluation scripts).
AMD ROCm Optimization
This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs. Key points:
Exported/tested with ROCm <rocm-version> and PyTorch ROCm build <torch-rocm-version>.
Validated backends: PyTorch (native ROCm HIP kernels) and vLLM (ROCm-enabled server build).
No code changes required versus the upstream OpenAI CLIP implementation — only environment/runtime configuration differs.
CPU fallback path supported for environments without a ROCm-capable GPU.
Runtime
Precision
Backend
Hardware
Notes
PyTorch
fp32/fp16
HIP (ROCm)
AMD Instinct / Radeon GPU
Native OpenAI CLIP inference
PyTorch
fp32
CPU
AMD CPU (EPYC/Ryzen)
CPU-only fallback
vLLM
fp16
ROCm server
AMD Instinct GPU
Pooling runner, image-only requests
vLLM (client)
—
HTTP
AMD CPU
Text-prompt requests, CPU client
Getting Started
Option 1: Use Provided Scripts
Pre-configured evaluation scripts are available for direct use on ROCm hardware. See Quick Start above for setup steps.
Option 2: Run with Custom Configuration
Use the scripts in on GitHub to run with your own:
Custom model variant (base16, large14, large14-336, etc.)
Custom dataset (beyond CIFAR-10)
Target AMD GPU/CPU and runtime (PyTorch vs vLLM)
This option is ideal if you need to customize the evaluation beyond the default configuration provided here.
Model Details
Model Type: Zero-shot image classification (contrastive image-text embedding)
Base Model:openai/clip-vit-base-patch32 (ViT-B/32)
Model Stats:
Model variant: base32 (default) — base16 / large14 / large14-336 also supported
Higher zero-shot accuracy means more test images are assigned the correct CIFAR-10 class via CLIP's image–text similarity — 100% is perfect, 10% is chance level for 10 classes. Values above ~85% on CIFAR-10 with ViT-B/32 are typical for this benchmark.
Metrics Explained
Metric
Description
Zero-shot accuracy (%)
Fraction of CIFAR-10 test images whose highest-scoring text prompt matches the ground-truth label after softmax over 10 class prompts. Primary accuracy metric; sensitive to both image and text embedding quality.
Accuracy Results
Full Dataset Evaluation (CIFAR-10 test) — filled from evaluation_results/; run make metrics to refresh:
Device
Backend
Precision
Variant
Accuracy (%)
CPU
PyTorch
FP16
clip-vit-base-patch32
88.79
CPU
PyTorch
FP32
clip-vit-base-patch32
88.80
GPU
PyTorch
FP16
clip-vit-base-patch32
88.75
GPU
PyTorch
FP32
clip-vit-base-patch32
88.80
GPU
vLLM
FP16
clip-vit-base-patch32
88.78
GPU
vLLM
FP32
clip-vit-base-patch32
88.80
Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?