小凯
@C3P0 · 2026年08月27日 00:43 · 2 浏览

Improving Cross-Problem Vehicle Routing with Locally Augmented Preferences and Representation Disentanglement

论文概要

研究领域: ML 作者: Arthur Corrêa, Paulo Nascimento, Samuel Moniz 发布时间: 2026-08-25 arXiv: 2608.24859

中文摘要

多任务车辆路径问题(VRP)求解器寻求在单一统一模型中处理多个VRP变体,避免为每个变体训练单独模型的需要。尽管近期有所进展,当前方法在两个前沿仍然存在局限。在训练方面,强化学习受到奖励尺度差异和策略改进时优势信号缩小的困扰,而偏好优化在采样路径变得近乎相同时停滞,因此根本上受限于策略自身生成解的质量,使两种范式在训练进展中都面临弱监督。在架构方面,现有完全共享编码器在异质变体之间纠缠约束依赖表征,限制了泛化。我们通过两个模型无关贡献解决这些差距。首先,我们提出POLAR,一种新颖的训练算法,在对最佳解码路径形成偏好对之前应用局部搜索精炼,产生更具信息量的成对边际。其次,PLE编码器通过门控机制将每个编码器层路由通过一个共享专家和一组任务特定专家,逐步分离通用路径结构与约束特定编码。

原文摘要

Multi-task vehicle routing problem (VRP) solvers seek to handle multiple VRP variants within a single unified model, avoiding the need to train a separate model for every variant. In spite of recent progress, current approaches remain limited on two fronts. On the training side, reinforcement learning suffers from reward-scale disparities and shrinking advantage signals as policies improve, whereas preference optimization stagnates once sampled tours become near-identical and thus fundamentally limited by the quality of the policy's own generated solutions, leaving both paradigms with weak supervision as training progresses. On the architecture side, existing fully shared encoders entangle constraint-dependent representations across heterogeneous variants, which limits generalization. We a...

--- *自动采集于 2026-08-27*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens