论文概要
研究领域: CV
作者: Olivier Dietrich, Krishna Sapkota, Konrad Schindler, Genady Beryozkin
发布时间: 2026-08-28
arXiv: 2608.28567
中文摘要
传统上,建筑物损伤评估(BDA)要么通过专用网络架构解决,要么通过微调地理空间图像基础模型解决。在这项工作中,我们探讨通用视觉-语言模型(VLM)是否能够仅通过自回归序列生成本地化建筑物并对其损伤进行分级。我们将BDA转换为预测可变长度的边界框集合,每个边界框由其坐标和损伤标签指定。我们的初步实现基于开放的Gemma模型,仅从双时相卫星图像和合适的文本提示就实现了有希望的损伤映射结果。
原文摘要
Conventionally, Building Damage Assessment (BDA) is tackled either with dedicated network architectures or by fine-tuning geospatial image foundation models. In this work, we ask whether a general-purpose Vision-Language Model (VLM) can localize buildings and grade their damage through autoregressive sequence generation alone. We cast BDA as predicting a variable-length set of bounding boxes, each specified by its coordinates and a damage label. Our preliminary implementation, based on the open Gemma model, achieves promising damage mapping results from only bi-temporal satellite images and a suitable text prompt.
自动采集于 2026-09-01
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。