[论文] Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
研究领域: NLP 作者: Tian Xia, Minghao Liu, Yiqing Liang, et al. 发布时间: 2026-09-30 arXiv: 2609.26730
论文概要
研究领域: NLP 作者: Tian Xia, Minghao Liu, Yiqing Liang, et al. 发布时间: 2026-09-30 arXiv: 2609.26730
中文摘要
多模态大语言模型(MLLM)正在快速推进临床诊断,但其适配流程仍锚定在基于准确率的目标上。临床数据普遍存在严重的类别不平衡:一个恒定预测多数类别的模型可以达到 90% 以上的准确率,却在临床上毫无价值。因此,我们针对 AUROC 进行评估和优化——这是一个无需阈值的评分指标,能将正样本排在负样本之上,且不受类别平衡影响。我们聚焦于 MLLM 中的提示词优化。GEPA 等反思式方法使用一个二元评分矩阵:每行对应一个评估实例,每列对应一个候选提示词;单元格记录逐实例的正确性,因此列平均值即为准确率,驱动候选选择。我们引入配对级帕累托提示词进化(Ranking-PE),将每个正确性行替换为对(正样本、负样本)实例对的逐对排序行:若候选提示词将正样本评分排在配对负样本之上,则单元格为 1。此时列平均值等于经验 AUROC(由 Wilcoxon-Mann-Whitney 恒等式保证)。我们在提示词进化搜索读取的三个层面都进行了这一替换——决定帕累托优势的评分矩阵、提供给反思 LM 的逐例反馈、以及最终候选选择——且不增加额外的模型调用,也不引入代理损失。在 MIMIC 数据集的三个疾病上,基于准确率的提示词进化反而可能恶化排序效果;Ranking-PE 扭转了这一趋势,在微调的 Qwen3-VL-8B 上超越基于准确率的方案 +5.8 个 AUROC 百分点,在 MedGemma-4B 上超越 +16.2 个百分点。
原文摘要
Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each co...
*自动采集于 2026-10-02*
#论文 #arXiv #NLP #小凯