[论文] WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answer...

研究领域: ML 作者: Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy, Navin Goyal 发布时间: 2026-09-15 arXiv: 2609.12171

论文概要

研究领域: ML 作者: Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy, Navin Goyal 发布时间: 2026-09-15 arXiv: 2609.12171

中文摘要

企业环境为问答智能体带来了严峻挑战——这类系统通常依赖检索增强生成、深度研究(DR)及相关技术。挑战很大程度上来自企业数据的复杂性:信息往往分散在不断演化且可能相互冲突的邮件、聊天消息、文档与其他工件中。现有基准通常真实复杂度有限、回答篇幅短、查询不自然,因此无法捕捉企业环境的真实挑战。我们提出一条自动化流水线,用于生成反映真实职场场景的合成邮件数据集,以及长/短形式的问题和扎根于数据的黄金答案。我们的方法模拟跨越数月、涉及多达 25 名不同角色员工交互的企业项目,数据强调歧义性、信息分散性与自然产生的查询。为验证流水线,我们使用最新的前沿模型在数据集上评估了若干标准智能体基线。结果发现,每个数据集上所有查询的平均总分均低于 80%,表明改进空间巨大。这些发现说明企业部署仍有大量工作要做,并凸显了真实、高复杂度评估数据对开发更强大的企业级深度研究系统的重要性。

原文摘要

Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation, Deep Research (DR), and related techniques. Much of this challenge comes from the complexity of enterprise data: information is often spread across evolving and potentially conflict- ing emails, chat messages, documents, and other artifacts. Existing benchmarks typically have limited real-world complexity, short-form responses, and unnatural queries, so they often fail to capture the challenges of enterprise settings. In this work, we introduce an automated pipeline for generating synthetic datasets of emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data. Our method simulates long...


*自动采集于 2026-09-15*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens