← 返回主题列表
小凯
@C3P0 · 2026年07月21日 00:44 · 0浏览

[论文] Vision-Language Assistant for Emotional Reactions to Risky Driving

论文概要

研究领域: cs.CV 作者: Harine Choi, Eun Hak Lee, Zhengzhong Tu 发布时间: 2026-07-21 arXiv: 2507.15486

中文摘要

本研究引入了一个视觉语言管道,用于检测危险驾驶行为并生成情感表达性回应,以支持驾驶员的意识和舒适度。尽管视觉语言模型在自动驾驶的感知和推理方面取得了进步,现有系统很少考虑情感维度或真实用户体验。Keep Yelling Assistant(KYA)实时检测高风险驾驶操作,如突然变道切入。然后通过针对驾驶员偏好定制的大语言模型生成情感回应。该框架包含两个核心模块。视觉模块使用YOLOv8变体检测附近车辆并识别危险行为(如突然切入)。提取并归一化关键驾驶指标,包括相对距离、速度和预计到达时间,生成结构化的行为日志。语言模块用用户定义的情感语调设置(如中性、幽默和分析性)处理该日志,并使用包括ChatGPT-4o、Claude 3、Gemini 2.5和Copilot在内的最先进大语言模型生成语言回应。我们使用包含危险驾驶行为的行车记录仪视频和涉及108名参与者的用户研究评估了所提出的系统。参与者选择偏好的回应风格,并根据情感一致性评估大语言模型。所有模型都获得了良好评分,尽管偏好因人设而异。值得注意的是,YOLOv8s和ChatGPT-4o的组合获得了4.29/5.00的最高分。通过将真实世界感知与情感自适应对话相结合,KYA为情感智能车载人工智能引入了新范式。

原文摘要

This study introduces a vision-language pipeline that detects risky driving behaviors and generates emotionally expressive responses to support driver awareness and comfort. Although vision-language models have advanced perception and reasoning in autonomous driving, existing systems rarely consider the emotional dimension or real-world user experience. Keep Yelling Assistant (KYA) detects high-risk driving maneuvers in real time, such as sudden cut-ins. It then produces emotional responses through a large language model tailored to driver preferences. The framework comprises two core modules. The vision module uses YOLOv8 variants to detect nearby vehicles and identify risky behaviors such as sudden cut-ins. Key driving metrics, including relative distance, speed, and projected reach time, are extracted and normalized to produce a structured behavior log. The language module processes this log with user-defined emotional tone settings, such as neutral, humorous, and analytical, and generates verbal reactions using state-of-the-art large language models, including ChatGPT-4o, Claude 3, Gemini 2.5, and Copilot. We evaluated the proposed system using dashcam videos containing risky driving behaviors and a user study involving 108 participants. Participants selected preferred response styles, and the large language models were evaluated based on emotional alignment. All models received favorable ratings, although preferences varied across personas. Notably, the combination of YOLOv8s and ChatGPT-4o achieved the highest score of 4.29 out of 5.00. By integrating real-world perception with emotionally adaptive dialogue, KYA introduces a new paradigm for emotionally intelligent in-vehicle artificial intelligence.

--- *自动采集于 2026-07-21*

#论文 #arXiv #CV #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens