PuSH - Publikationsserver des Helmholtz Zentrums München

Huang, H.* ; Zanca, D.* ; Eskofier, B.M. ; Salin, E.*

General or Medical CLIP, Which One Shall We Choose?

In: (Artificial Intelligence in Healthcare). Berlin [u.a.]: Springer, 2026. 74 - 87 (Lect. Notes Comput. Sc. ; 16877 LNCS)
DOI
Despite growing interest in multimodal deep learning for medical imaging, researchers and clinicians still lack a systematic understanding of when medical-domain vision–language models outperform their general-domain counterparts. While medical CLIP variants have been proposed, they are typically evaluated in isolation and on narrow tasks, leaving open questions about how pre-training data, downstream task, and fine-tuning strategy jointly affect performance. We systematically compare four CLIP-based models on three representative tasks: image classification, image-to-text retrieval, and visual question answering. Across tasks, zero-shot performance is generally insufficient for clinical use, even for medically pre-trained models, confirming the need for task-specific fine-tuning. Medical-domain pre-training offers clear benefits in low-data regimes and for in-distribution modalities, but can underperform CLIP when downstream data deviates from the pre-training distribution. When sufficient labeled data is available, and especially under LoRA-based tuning, general-domain CLIP systematically matches or surpasses specialized medical models. VQA remains notably challenging, with none of the evaluated models achieving competitive results even after fine-tuning, suggesting that more advanced multimodal reasoning approaches are needed. Based on these findings, we provide recommendations for selecting and adapting CLIP-based models in clinical settings.
Altmetric
Weitere Metriken?
Zusatzinfos bearbeiten [➜Einloggen]
Publikationstyp Artikel: Konferenzbeitrag
Schlagwörter Downstream (manufacturing) ; Isolation (microbiology) ; Deep Learning ; Clinical Practice ; Medical Literature ; Exploit
ISSN (print) / ISBN 0302-9743
e-ISSN 1611-3349
Konferenztitel Artificial Intelligence in Healthcare
Quellenangaben Band: 16877 LNCS, Heft: , Seiten: 74 - 87 Artikelnummer: , Supplement: ,
Verlag Springer
Verlagsort Berlin [u.a.]