Research team introduces InfiMed2 medical multimodal models
The research paper introduces InfiMed2 – medical multimodal foundation models in versions with 4B and 27B parameters. According to the authors, the 4B model achieves 66.73% accuracy after RLVR and outperforms the larger Qwen3.5-9B, while the 27B model achieves 73.72 % on five benchmarks.
The research team introduced InfiMed2, a family of generalist medical multimodal foundation models in two sizes – 4B and 27B parameters. The models are built on a training corpus of 55.68 billion tokens that combines broad clinical knowledge with context-rich biomedical visual evidence processed according to source type.
Training proceeds in stages: first, adaptation of the visual encoder, then the development of broad medical knowledge, and finally a transition to an evidence-focused data mixture during the learning rate decay phase. For supervised fine-tuning, the team regenerated answers to visual questions using answer stability, answer-masked reconstruction and correctness-constrained selection to achieve more informative and consistent answers. The 4B model is further optimized using reinforcement learning with verifiable rewards (RLVR).
According to the authors, InfiMed2-4B achieves an average accuracy of 66.73 % across five medical multimodal benchmarks after RLVR, outperforming the larger model Qwen3.5-9B. The model InfiMed2-27B achieves an accuracy of 73.72 %, which, according to the paper, is the highest value among the evaluated open-weight models.
Why it matters
The results show that carefully designed training (stage-aware data processing, RLVR) may help a smaller model outperform larger competing models on medical tasks, which is relevant to developers of specialized healthcare AI applications considering the trade-off between model size and performance.
Two audiences, two different impacts
What this means
For individuals
Researchers and developers working on medical multimodal models gain a new reference benchmark and a description of an approach (stage-aware data design, answer-stability supervision, RLVR) that they can consider in their own development.
More practical updates →For a business
Companies developing medical AI tools can infer from the results that a smaller specialized model (4B parameters) may outperform larger general-purpose models with a suitable training approach, which may reduce computational requirements during deployment.
Development More business impacts →Check the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.