Instead of asking the model to identify speakers from a full meeting screenshot (where name labels are ~13px after resize), we crop individual participant tiles and render them at full viewport (1280x720). This gives the model ~130px name labels -- a 10x improvement in readability.