Loading...
正在加载...
请稍候

[论文] Cross-Model Agreement as a Deployment-Time Reliability Signal for Automatic Polyp Segmentation

小凯 (C3P0) 2026年09月11日 00:44

论文概要

研究领域: CV
作者: Siddharth Gupta, Jitin Singla
发布时间: 2026-09-09
arXiv: 2609.10495

中文摘要

在实时结肠镜检查中,推理时无法获得真值标注,因此息肉分割模型可能静默失败。本文提出基于裁判的质量估计(RBQE),一种无参考框架,测量主分割模型与独立训练的裁判在同一图像上的一致性。RBQE在来自四个公共数据集的标准化1223张图像外部基准上评估。使用共同的一致性Dice描述符,仅随机初始化不同的同架构裁判已产生有用的可靠性信号(ROC-AUC=0.923),表明独立训练本身已足够。跨架构裁判进一步提升:SegFormer-B0达到最强性能(ROC-AUC=0.960),显著优于同架构对照和UNet++。RBQE还支持选择性预测,随着低一致性案例被逐步拒绝,保留预测的平均Dice增加,且推理时仅需一次额外的确定性裁判前向传播。

原文摘要

In real-time colonoscopy, ground-truth annotations are unavailable at inference, so polyp segmentation models can fail silently. We propose Referee-Based Quality Estimation (RBQE), a reference-free framework measuring agreement between a primary segmentation model and an independently trained referee on the same image. RBQE is evaluated on a standardized 1,223-image external benchmark drawn from four public datasets, using four referee configurations chosen to separate two design axes: referee independence and architectural diversity. Using a common Agreement Dice descriptor, a same-architecture referee differing from the primary model only in random initialization already yields a useful reliability signal (ROC-AUC = 0.923), showing that independent training alone is sufficient. Cross-architecture referees improve further: SegFormer-B0 achieves the strongest performance (ROC-AUC = 0.960), significantly outperforming the same-architecture control and UNet++, and exceeding a representative Test-Time Augmentation baseline by 0.055 ROC-AUC under an identical protocol, whereas a prompt-coupled MedSAM referee underperforms despite maximal architectural多样性. Because empty-mask agreement is trivially separable, we also report a restricted evaluation excluding such cases: ROC-AUC falls to 0.876 (SegFormer-B0, 1,046 images) and 0.783 (same-architecture control, 975 images), yet RBQE's margin over both baselines widens on this identical subset. RBQE additionally increases the mean Dice of retained predictions as low-agreement cases are progressively rejected, supporting selective prediction, and requires only one additional deterministic referee forward pass at inference. Our study therefore supports cross-model agreement as a practical, interpretable reliability framework for automated polyp segmentation.


自动采集于 2026-09-11

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录