论文概要
研究领域: ML
作者: Akshay Balsubramani
发布时间: 2026-08-18
arXiv: 2608.18061
中文摘要
我们提出了一个学习者和自然之间的两人零和重复博弈,其价值恒等式同时生成了贝叶斯更新和指数权重后悔的精确核算,并提供了一类广泛集中现象共享的比较器类变分形式。终局收益是比较器在相对于先验的固定相对熵下能获得的最大增益,而单步约束是学习者混合行动下自然移动的信息预算。在学习者的行动不受其他限制的情况下,Gibbs/贝叶斯权重作为其唯一的Bellman均衡器出现——这种混合行动使每轮损失独立于自然移动的方向——对数分区函数扮演价值函数的角色。后悔精确分解为三部分:反映观测结果变化的每轮信息损失、精确解释各轮之间测量尺度变化的加性再回火漂移,以及比较器相对于先验携带的信息。驱动标准后悔边界的方差和有界范围代理是这种分解的更宽松的松弛,而该分解普遍成立并支配着它们。两个玩家的策略从分解中逐项读出,重复博弈产生了自博弈的信息论账本,代替了通常的二次变分替代。相同的比较器类几何解释了经典的大偏差边界,而bandits、后验采样、聚合和提升中的方法都是单一后悔分解的特例。
原文摘要
We give a two-player zero-sum repeated game between a learner and nature whose value identity generates Bayesian updating and an exact accounting of exponential-weights regret at once, and supplies the comparator-class variational form that a wide class of concentration phenomena share. The terminal payoff is the most a comparator can gain at fixed relative entropy from the prior, and the one-step constraint is an information budget on nature's move under the learner's mixed action. With the learner's move otherwise unrestricted, Gibbs/Bayes weights emerge as its unique Bellman equalizer -- the mixed action that makes the per-round loss independent of which direction nature moves -- with log-partition functions playing the role of value functions. The regret decomposes exactly into three p...
自动采集于 2026-08-20
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。