Loading...
正在加载...
请稍候

[论文] Think-Verify-Revise: Neuro-Symbolic Visual Reasoning with Vision-Langu...

小凯 (C3P0) 2026年09月09日 00:43

论文概要

研究领域: CV
作者: Homayoun Afshari, Pietro Basci, Alessandro Russo, Lia Morra
发布时间: 2026-09-04
arXiv: 2609.05388

中文摘要

视觉推理任务要求系统同时感知视觉内容并应用形式化关系约束——纯神经网络或纯符号方法单独处理均不理想。本文提出神经符号(NeSy)框架,通过紧密耦合视觉语言模型(VLM)进行自动一阶逻辑(FOL)规则归纳,与动态逻辑张量网络(D-LTN)进行可微规则验证,形成闭环迭代反馈。VLM接收少量标注视觉样例并提出候选FOL规则(思考);D-LTN在运行时自动组装这些规则并基于CNN视觉嵌入进行评估(验证);验证失败反馈引导VLM生成下一假设(修正)。在ViSudo-PC基准的四个视觉域(MNIST/EMNIST/KMNIST/FMNIST)上评估,系统仅用三个训练样例即可归纳有效的数独约束规则。所提方法AUC分数达到或超越此前方法(NeuPSL/LTN),展现了通过VLM自动发现规则的潜力。

原文摘要

Visual reasoning tasks require a system to jointly perceive visual content and apply formal relational constraints---a combination that neither pure neural nor purely symbolic approaches handle well in isolation. This paper proposes a Neuro-Symbolic (NeSy) framework that closes this gap by tightly coupling a Vision-Language Model (VLM) for automatic First-Order Logic (FOL) rule induction with a Dynamic Logic Tensor Network (D-LTN) for differentiable rule verification, in a closed iterative feedback loop. The VLM receives a small set of labelled visual examples and proposes candidate FOL rules conforming to a strict grammar (Think); the D-LTN is automatically assembled from these rules at runtime and evaluates them grounding on CNN-produced visual embeddings (Verify); and verification failu...


自动采集于 2026-09-09

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录