eval.py) is designed to benchmark over a dozen state-of-the-art Vision-Language Models (VLMs) against the SolarVQA dataset, automatically extracting answers, calculating chance-corrected metrics, and generating comprehensive leaderboards.transformers and specific backend wrappers:1pip install torch torchvision transformers accelerate
2pip install Pillow numpy pandas scikit-learn tqdm