"Can a $35 computer see, think, and chase a drone — all at once?"
That was the question. This is the answer.
🎯 The Challenge
Drones are everywhere. Detecting them in real-time on powerful hardware is one thing — but what if you only have a Raspberry Pi 4? No GPU, no CUDA, no shortcuts.
The challenge was clear:
Detect drones accurately in real-time
Run inference at usable frame rates on a CPU
Move a physical camera to follow the drone — automatically
Most people would say it's not possible. We built it anyway.
📊 Final Model Performance
Metric
Score
Precision
94.8%
Recall
96.2%
mAP@0.5
99.0% 🔥
mAP@0.5:0.95
70.6%
False Positives
~1% ✅
Inference on Pi
~15-20 FPS
💡 The Story Behind This Model
Act 1 — The First Attempt
We started simple. A single dataset, ~500 images, YOLOv5s with Transfer Learning. The first model came out at 95% mAP — impressive on paper.
But on the Raspberry Pi, reality hit hard.
The camera would lock onto a wooden tray and call it a drone. A chair. A lamp. The model was too eager — it had never truly learned what wasn't a drone.
We had a precision problem. And we knew exactly why: not enough negative images.
Act 2 — The Dataset Problem
The original dataset had only 5 background images out of 503 total. Five. The model had barely seen what the world looks like without a drone in it.
So we went hunting for data.
Instead of spending weeks manually collecting and labeling images, we combined 3 of the largest drone datasets on Roboflow Universe:
Dataset
Images
Source
DroneDetectionPITT
34,000
Roboflow Universe
Drone Detection (YOLO)
1,094
Roboflow Universe
Drone Detection (drones-lfobz)
5,079
Roboflow Universe
Positive (auto-labeled)
321
Original dataset
Negative (background)
155
COCO Dataset
Total
40,529
—
But combining datasets isn't just copy-paste. Each dataset has its own format, resolution, and labeling style. We built a custom pipeline to:
Download all 3 datasets via Roboflow API
Convert every label to unified YOLO format
Auto-label our positive images using the first model as a labeling tool
Add COCO background images as hard negatives
It took engineering. It took iteration. But it worked.
Act 3 — Training at Scale
With 40,000+ images ready, we needed serious compute. We upgraded to an NVIDIA A100 40GB GPU on Google Colab and trained for 30 epochs with:
Model : YOLOv5s (Transfer Learning from COCO)
Image size: 416×416
Batch size: 256
Epochs : 30
GPU : NVIDIA A100 40GB
Time : ~45 minutes
The results were immediate. mAP jumped from 95% to 99%. False positives dropped from 100% to 1%. The model had finally learned the difference between a drone and everything else.
Act 4 — Running on the Edge
A great model means nothing if it can't run where it needs to. We exported to ONNX format and deployed on Raspberry Pi 4 using ONNXRuntime — no GPU, no PyTorch, just pure optimized inference.
Achieved: 15-20 FPS on a $35 computer. ✅
Act 5 — The Tracker
Detection alone wasn't enough. We built a complete pan-tilt tracking system: