TransFPN-YOLO is a custom object detection model built on YOLOv8n with three targeted architectural modifications for crop stress detection in real-world field images.
Mean width: 0.088 · Mean height: 0.116 · Mean area: 0.016
The right-skewed area distribution confirms most objects are small, justifying the 1280px input upgrade.
66.2% of bounding boxes are small (<1% image area).
At 640px input → small object edge ≈ 64px. At 1280px → 128px. Doubling resolution makes small objects reliably detectable, at the cost of batch size (16→4) within T4 VRAM budget.
Architecture
Architecture Diagram
Modifications over Standard YOLOv8n
Component
YOLOv8n Baseline
TransFPN-YOLO (Ours)
Backbone Attention
None
CBAM after P3 & P4 CSP stages
Bottleneck
Standard C2f
C2f + Transformer Encoder (MHA) at P5
Neck
Standard PANet FPN
Weighted BiFPN (bidirectional)
Classification Loss
BCE
Focal Loss (class imbalance)
Parameters
~3.2M
~4.8M (+50%)
Input Size
640×640
1280×1280
Batch Size (T4)
16
4 (AMP-safe)
CBAM focuses the backbone on discriminative stress features via channel + spatial attention. Transformer Bottleneck captures long-range dependencies — useful for distributed disease lesion patterns. BiFPN fuses features bidirectionally with learned scalar weights — directly improves small object detection. Focal Loss down-weights easy majority-class examples, focusing training on the rare abiotic class.
The baseline outperforms on mAP/recall in this 50-epoch run because TransFPN-YOLO at batch=4 and 1280px needs more epochs to converge. However, TransFPN-YOLO achieves better mean IoU (0.583 vs 0.516) — boxes it predicts are more accurately localized, confirming the BiFPN neck improves localization quality.
Precision-Recall Curves
PR Curves
Class
Baseline AP
TransFPN-YOLO AP
Abiotic
0.925
0.693
Insect
0.893
0.687
Disease
0.969
0.920
Per-Class F1 Scores
F1 Scores
Per-Class AP@50
Per-Class AP
Class
Baseline
TransFPN-YOLO
Δ
Abiotic
0.221
0.066
−0.154
Insect
0.163
0.083
−0.080
Disease
0.121
0.039
−0.082
Confusion Matrices
Confusion Matrices
Notably, TransFPN-YOLO improves disease classification accuracy (0.76 vs baseline 0.69), consistent with the Transformer bottleneck and BiFPN helping detect distributed lesion texture patterns.
Radar Chart — Overall Performance Profile
Radar Chart
IoU Distribution
IoU Distribution
TransFPN-YOLO shows far fewer near-zero IoU detections (3 vs 45 in the 0–0.05 bin) and concentrates matched predictions at higher IoU values (0.6–0.85), confirming better localization quality per prediction.
Training Curves
Training Curves
Baseline converges more stably across 50 epochs. TransFPN-YOLO starts with higher box loss (expected for 1280px heavier architecture) and shows more val variance at batch=4 — indicating more training epochs are needed to fully leverage the added capacity.
Qualitative Comparison
Qualitative Results
Baseline detects more objects overall (higher recall). TransFPN-YOLO produces fewer but more precisely localized detections.