🏀 SportsMOT Visual Recommendation System
📌 Overview
This project implements a visual recommendation system based on image and text embeddings, built as part of Assignment #3 – Embeddings, Recommendation Systems, and Spaces.
The system allows users to:
• Upload an image and receive visually similar images
• Enter a textual description and retrieve matching images
• Explore similarities within a large-scale sports video dataset
The application is deployed as an interactive Gradio app on HuggingFace Spaces.
⸻
📂 Dataset
• Name: SportsMOT
• Source: HuggingFace Datasets
• Domain: Sports video frames (e.g., soccer, basketball, gameplay scenes)
Due to the dataset’s size, a balanced subset of images was used to ensure:
• Computational efficiency
• Preservation of visual structure and diversity
⸻
🧠 Model & Embeddings
• Model: OpenAI CLIP (ViT-B/32)
• Embedding Dimension: 512
• Modalities Supported: Image & Text
Each image in the dataset was converted into a normalized embedding vector.
All embeddings were saved to disk for reuse.
Stored File:
• clip_embeddings.parquet
Contains:
• Image embeddings
• Corresponding image paths
⸻
📊 Embeddings Analysis & Clustering
To analyze the embedding space:
• Dimensionality Reduction: PCA, UMAP
• Clustering Algorithm: K-Means
The clusters showed semantic coherence, grouping images by:
• Player close-ups
• Field-wide scenes
• Crowd and background-heavy frames
⸻
🔍 Recommendation Pipeline
1. User provides image or text
2. Input is converted into an embedding using CLIP
3. Cosine similarity is computed against dataset embeddings
4. Top-K most similar images are returned
⸻
🖥️ Application (Gradio Interface)
The app provides two input modes:
• Image Input: Upload an image and receive similar frames
• Text Input: Describe a scene and retrieve matching images
Users can control the number of returned results using a Top-K slider.
⸻
🚀 Deployment
• Platform: HuggingFace Spaces
• Framework: Gradio
• Language: Python
All required files (app.py, requirements.txt, and embeddings file) are included in the Space root directory.
🎥 Demo Video
A short presentation video (3–5 minutes) demonstrates:
• Dataset overview
• Embedding generation
• Clustering results
• Live usage of the deployed application
⸻
✅ Key Takeaways
• Demonstrates practical use of multimodal embeddings
• Combines computer vision with recommendation systems
• Fully deployed and reproducible interactive app