Multimodal Large Language Models in Medical Applications
Multimodal Large Language Models (MLLMs) have demonstrated potential in addressing medical applications, such as Medical Visual Question Answering (MVQA) and Medical Report Generation (MRG), by integrating visual and textual data. This research proposes a novel fine-tuning framework that will leverage a vision encoder, a pre-trained Large Language Model (LLM), and a lightweight linear projection layer to align visual embeddings with the LLM space. The proposed framework will employ a multimodal prompt template that incorporates visual features, textual questions, and task-specific identifiers to enable the LLM to generate accurate, context-aware outputs. To optimize resource efficiency, Low-Rank Adaptation (LoRA) will be utilized for fine-tuning, while other components of the model will remain frozen. A two-stage fine-tuning strategy will be developed to ensure scalability and robustness for various medical tasks. To evaluate the performance of the proposed framework, we will design an evaluation approach for assessing the quality of generated outputs, complemented by traditional metrics. Comprehensive validation will be conducted using publicly available benchmark datasets for both MVQA and MRG tasks. This research aims to enhance the efficiency, scalability, and clinical applicability of MLLMs in medical domains while addressing gaps in existing methodologies