Views
No views yet
| Architecture | ViT | LLM | Adapter | Resolution |
|---|---|---|---|---|
| 🤗InfiMed-Foundation-4B | 🤗siglip-so400m-patch14-384 | 🤗Qwen3-4B | 2-layer MLP | 384x384xN |
| Model | Size | MMMU-Med | VQA-RAD | SLAKE | PathVQA | PMC-VQA | OMVQA | MedXVQA | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| Proprietary Models | |||||||||
| GPT-5 | 83.6 | 67.8 | 78.1 | 52.8 | 60.0 | 76.4 | 71.0 | 70.0 | |
| GPT-5-mini | 80.5 | 66.3 | 76.1 | 52.4 | 57.6 | 70.9 | 60.1 | 66.3 | |
| GPT-5-nano | 74.1 | 55.4 | 69.3 | 45.4 | 51.3 | 66.5 | 45.1 | 58.2 | |
| GPT-4.1 | 75.2 | 65.0 | 72.2 | 55.5 | 55.2 | 75.5 | 45.2 | 63.4 | |
| Claude Sonnet 4 | 74.6 | 67.6 | 70.6 | 54.2 | 54.4 | 65.5 | 43.3 | 61.5 | |
| Gemini-2.5-Flash | 76.9 | 68.5 | 75.8 | 55.4 | 55.4 | 71.0 | 52.8 | 65.1 | |
| General Open-source Models | |||||||||
| Qwen2.5VL-3B | 3B | 51.3 | 56.8 | 63.2 | 37.1 | 50.6 | 64.5 | 20.7 | 49.2 |
| Qwen2.5VL-7B | 7B | 54.0 | 65.0 | 67.6 | 44.6 | 51.3 | 63.5 | 21.7 | 52.5 |
| InternVL3-8B | 8B | 59.2 | 65.4 | 72.8 | 48.6 | 53.8 | 79.1 | 22.4 | 57.3 |
| Medical Open-source Models | |||||||||
| MedGemma-4B-IT | 4B | 43.7 | 72.5 | 76.4 | 48.8 | 49.9 | 69.8 | 22.3 | 54.3 |
| LLaVA-Med-7B | 7B | 29.3 | 53.7 | 48.0 | 38.8 | 30.5 | 44.3 | 20.3 | 37.8 |
| HuatuoGPT-V-7B | 7B | 47.3 | 67.0 | 67.8 | 48.0 | 53.3 | 74.2 | 21.6 | 54.2 |
| Lingshu-7B | 7B | 54.0 | 67.9 | 83.1 | 61.9 | 56.3 | 82.9 | 26.7 | 61.8 |
| BioMediX2-8B | 8B | 39.8 | 49.2 | 57.7 | 37.0 | 43.5 | 63.3 | 21.8 | 44.6 |
| Infi-Med-1.7B | 1.7B | 34.7 | 56.3 | 75.3 | 60.7 | 48.1 | 58.9 | 21.8 | 50.8 |
| Infi-Med-4B | 4B | 43.3 | 57.9 | 77.7 | 63.4 | 56.6 | 76.8 | 21.9 | 56.4 |
1git clone https://huggingface.co/InfiX-ai/InfiMed-Foundation-4B
2cd InfiMed-Foundation-4B1from InfiMed import InfiMed
2from PIL import Image
3import torch
4
5# Load the model from the pretrained checkpoint
6model = InfiMed.from_pretrained("InfiX-ai/InfiMed-Foundation-4B", device_map="auto", torch_dtype=torch.bfloat16)
7
8image_path = "sample.png" # Replace with the path to your image file
9image = Image.open(image_path).convert("RGB") # Ensure the image is in RGB format
10
11# Prepare input messages
12messages = {
13 "prompt": "What modality is used to take this image?",
14 "image": image # No image for this example, set to None
15}
16
17# Generate output
18output_text = model.generate_output(messages)
19
20# Print the result
21print("Model Response:", output_text)
22