Views
No views yet
lm_head.pt -- Preserved lm_head weights from Qwen2.5-Omni-3B thinkerresults/ -- Raw evaluation JSONsscripts/ -- Training + eval scripts| Benchmark | avg nDCG@5 | # tasks |
|---|---|---|
| ViDoRe V1 | 0.8865 | 10 |
| ViDoRe V2 | 0.5353 | 4 |
| ViDoRe V3 | 0.4907 | 8 |
| AudioCaps R@1 (zero-shot) | 26.2% | — |
results/VidoreInfoVQARetrieval_predictions.json), directly comparable to the HydraQwen3.5-4B V1 column.| n | Base ANLS | Hydra ANLS | Δ (95% CI) | Exact match |
|---|---|---|---|---|
| 2,801 | 0.7257 | 0.7257 | +0.0000 [+0.0000, +0.0000] | 2,801 / 2,801 (100.00%) |
results/infovqa_report.json.1@article{georgiou2026hydra,
2 title={Hydra: Unifying Document Retrieval and Generation in a Single Vision-Language Model},
3 author={Georgiou, Athos},
4 year={2026}
5}