论文概要
研究领域: cs.AI
作者: EunKyeong Lee, Kyeong-Jin Oh, Jinwon Kim, Hye Woo Lee, Minsang Song, Hyeongjun Jang, Junyoung Youn
发布时间: 2026-09-13
arXiv: 2609.11065
中文摘要
图检索增强生成(GraphRAG)可以连接跨语料图分布的证据,但大多数系统在各查询间使用基本共享的探索过程。这造成结构不匹配:直接事实可能需要紧凑的本地邻域,比较需要多个目标的平衡覆盖,中介问题可能需要通过弱相关连接器的更深路径。我们提出 Mosaic,一个免训练框架,将 GraphRAG 检索形式化为每查询控制问题。LLM 分析器将查询特定的证据需求转换为种子选择、图遍历、停止和证据选择的有界策略,而语料图、索引、评分函数、接地过程和答案生成器保持共享。在 GraphRAG-Bench 上,Mosaic 在 Medical 上实现查询加权答案正确性 76.97,在 Novel 上 64.33,比之前报告的最强总体结果提高 5.13 和 4.43 点。在 Medical 上达到 95.1 证据召回率和 86.1 上下文相关性。相对于固定宽策略,它评估少 81.9% 的路径并保留少 47.2% 的证据项。
原文摘要
Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, comparisons need balanced coverage of multiple targets, and mediated questions may require deeper paths through weakly related connectors. We present Mosaic, a training-free framework that formulates GraphRAG retrieval as a per-query control problem. An LLM analyzer converts query-specific evidence requirements into a bounded policy over seed selection, graph traversal, stopping, and evidence selection, while the corpus graph, indexes, scoring functions, grounding procedure, and answer generator remain shared.
On GraphRAG-Bench, Mosaic achieves query-weighted Answer Correctness of 76.97 on Medical and 64.33 on Novel, improving over the strongest previously reported overall results by 5.13 and 4.43 points. On Medical, it reaches 95.1 Evidence Recall and 86.1 Context Relevancy. Controlled comparisons on an identical graph and generator show that no fixed narrow, medium, or wide policy is consistently optimal; Mosaic improves by 9.96 points over the strongest canonical fixed policy. Relative to Fixed Wide, it evaluates 81.9% fewer paths and retains 47.2% fewer evidence items. Transfer experiments on HotpotQA, MuSiQue, and 2WikiMultiHopQA further show that the policy interface can be applied without benchmark-specific retriever training.
自动采集于 2026-09-13
#论文 #arXiv #AI #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。