Views
No views yet
| Component | Details |
|---|---|
| Text Encoder | XLM-RoBERTa base (frozen) |
| Image Encoder | EVA02-CLIP-L/14 from QuanSun (frozen) |
| Feature Dim | 768 projected to 512 |
| Fusion | Dual Cross-Attention + Adaptive Gating |
| Conflict Detection | Cosine similarity + MLP |
| Reliability Weighting | Per-modality sigmoid scoring |
| Classifier | MLP (512 -> 256 -> 128 -> 3) |
| Split | Ratio |
|---|---|
| Train | 70% |
| Validation | 15% |
| Test | 15% |
| Parameter | Value |
|---|---|
| Epochs | 50 (early stopped at 26) |
| Optimizer | AdamW (lr=1e-4) |
| Scheduler | CosineAnnealingLR |
| Loss | Focal Loss + Class Weights |
| Early Stopping | Patience = 7 |
| Batch Size | 32 |
| Max Text Length | 128 tokens |
| Image Size | 224 x 224 |
| GPU | T4 (Kaggle) |
| Metric | Score |
|---|---|
| Macro-F1 | 0.5584 |
Full test set metrics will be updated after evaluation.
Main metric: Macro-F1
