English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GeBDA: Building Damage Assessment as Text-Based Sequence Prediction with Vision-Language Models

Forum topic · 小凯 · 2026-09-01

Summary

GeBDA (arXiv:2608.28567) explores whether general-purpose Vision-Language Models can perform Building Damage Assessment (BDA) purely through autoregressive sequence generation, without dedicated network architectures or fine-tuned geospatial foundation models. The authors cast BDA as predicting a variable-length set of bounding boxes, where each box is specified by its coordinates and a damage label, all generated as text. A preliminary implementation built on the open Gemma model achieves promising damage mapping results using only bi-temporal satellite images and a suitable text prompt. By Olivier Dietrich, Krishna Sapkota, Konrad Schindler, and Genady Beryozkin, the work suggests that text-based sequence prediction offers a flexible, prompt-driven alternative for geospatial damage mapping tasks.

Paper Overview

Research area: Computer Vision Authors: Olivier Dietrich, Krishna Sapkota, Konrad Schindler, Genady Beryozkin arXiv: 2608.28567

Abstract

Conventionally, Building Damage Assessment (BDA) is tackled either with dedicated network architectures or by fine-tuning geospatial image foundation models. In this work, the authors ask whether a general-purpose Vision-Language Model (VLM) can localize buildings and grade their damage through autoregressive sequence generation alone. BDA is cast as predicting a variable-length set of bounding boxes, each specified by its coordinates and a damage label. A preliminary implementation, based on the open Gemma model, achieves promising damage mapping results from only bi-temporal satellite images and a suitable text prompt.

Key Points

  • Replaces specialized BDA pipelines with a general-purpose VLM using text-based autoregressive generation.
  • Task formulation: predict a variable-length set of bounding boxes as text, each with coordinates and a damage label.
  • Prototype built on the open Gemma model.
  • Input requires only bi-temporal satellite images plus a text prompt — no task-specific fine-tuning in the preliminary version.
Source: arXiv:2608.28567

Tags

#computer-vision#vision-language-models#building-damage-assessment#satellite-imagery#gemma#geospatial#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634335