Summary
This arXiv paper (2506.17580) by Gina Wong, Drew Prinster, and Suchi Saria studies calibration in Mixture-of-Experts (MoE) models under distribution shift. Calibration aligns a model's predictive uncertainty with empirical outcome frequencies, which is essential for interpreting and trusting reported probabilities. Prior work has shown that enforcing calibration at the level of individual predictors can improve the accuracy and calibration of ensemble models, with MoE showing particularly notable empirical gains—but the conditions under which calibration benefits MoE remained unclear. The authors analyze MoE behavior under distribution shift, focusing on the interaction between routing mechanisms and expert-level calibration. They prove that expert calibration is sufficient to guarantee overall model calibration for hard-routed models under a broad class of distribution shifts, but is insufficient for soft-routed models. To address this, they propose an adversarial reweighting method that penalizes the calibration error of the routing aggregate under shift, demonstrating improved accuracy–calibration trade-offs both on average and on hard data subsets, across model classes, prediction tasks, and types of distribution shift.
Paper Overview
- Research area: cs.AI, cs.LG
- Authors: Gina Wong, Drew Prinster, Suchi Saria
- Published: 2026-06-21
- arXiv: 2506.17580
Summary
Calibration — aligning a model's predictive uncertainty with the empirical frequency of outcomes — is essential for interpreting and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can improve the accuracy and calibration of ensemble models, with Mixture-of-Experts (MoE) models exhibiting particularly notable empirical gains. However, the conditions under which calibration benefits MoE have remained unclear.
This work studies MoE model behavior under distribution shift, focusing on the interaction between routing mechanisms and expert-level calibration. The authors prove that expert calibration is sufficient to ensure overall model calibration for hard-routed models under a broad class of distribution shifts, but it is insufficient for soft-routed models.
To address this gap, the paper proposes an adversarial reweighting method that penalizes the calibration error of the routing aggregate under distribution shift. The method is shown to improve the accuracy–calibration trade-off, both on average and on hard subsets of the data, across model classes, prediction tasks, and distribution shifts.
---
*Automatically collected on 2026-06-21*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178207994