Views
No views yet
CRAI_composite = 0.30*CEA + 0.20*CC + 0.20*CS + 0.20*CI - 0.10*HP| system | Spearman ↑ | MAE ↓ |
|---|---|---|
| Organiser GPT-4 baseline (as reported) | 0.6250 | 0.2248 |
GPT-4 raw, scoring the shipped gold_llm.tsv | 0.2153 | 0.3666 |
| Qwen3-VL-8B raw | 0.5482 | 0.2550 |
| Qwen3-VL-8B calibrated + version | 0.7064 | 0.1960 |
| version one-hot only — no image, no judge | 0.6823 | 0.2091 |
[5 raw dims + caption-version one-hot] to the human scores,
fit on the 120 train instances. The composite is always recomputed with the official
formula, never predicted.gold_llm.tsv
scores 0.2153 on dev, not the reported 0.6250.