论文概要
研究领域: CV
作者: Homayoun Afshari, Pietro Basci, Alessandro Russo, Lia Morra
发布时间: 2026-09-04
arXiv: 2609.05388
中文摘要
视觉推理任务要求系统同时感知视觉内容并应用形式化关系约束——纯神经网络或纯符号方法单独处理均不理想。本文提出神经符号(NeSy)框架,通过紧密耦合视觉语言模型(VLM)进行自动一阶逻辑(FOL)规则归纳,与动态逻辑张量网络(D-LTN)进行可微规则验证,形成闭环迭代反馈。VLM接收少量标注视觉样例并提出候选FOL规则(思考);D-LTN在运行时自动组装这些规则并基于CNN视觉嵌入进行评估(验证);验证失败反馈引导VLM生成下一假设(修正)。在ViSudo-PC基准的四个视觉域(MNIST/EMNIST/KMNIST/FMNIST)上评估,系统仅用三个训练样例即可归纳有效的数独约束规则。所提方法AUC分数达到或超越此前方法(NeuPSL/LTN),展现了通过VLM自动发现规则的潜力。
原文摘要
Visual reasoning tasks require a system to jointly perceive visual content and apply formal relational constraints---a combination that neither pure neural nor purely symbolic approaches handle well in isolation. This paper proposes a Neuro-Symbolic (NeSy) framework that closes this gap by tightly coupling a Vision-Language Model (VLM) for automatic First-Order Logic (FOL) rule induction with a Dynamic Logic Tensor Network (D-LTN) for differentiable rule verification, in a closed iterative feedback loop. The VLM receives a small set of labelled visual examples and proposes candidate FOL rules conforming to a strict grammar (Think); the D-LTN is automatically assembled from these rules at runtime and evaluates them grounding on CNN-produced visual embeddings (Verify); and verification failu...
自动采集于 2026-09-09
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。