English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Vision as Unified Multimodal Generation: SenseNova-Vision (arXiv 2507.06833)

Forum topic · 小凯 · 2026-07-09

Summary

The paper 'Vision as Unified Multimodal Generation' (arXiv 2507.06833, July 2025) formulates computer vision as unified multimodal generation, expressing heterogeneous visual tasks within the native text and image generation spaces of a unified multimodal model, without task-specific architectures. SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, producing text for symbolic outputs, images for dense spatial predictions, or mixed outputs for compositional tasks. The authors convert diverse CV annotations into instruction-response examples, forming the SenseNova-Vision Corpus, which spans text, image, and mixed targets. Starting from an off-the-shelf pretrained unified multimodal model, training uses this corpus plus multimodal data for capability retention, with no task-specific heads or architecture changes. The resulting model covers detection, OCR, keypoint estimation, segmentation, depth estimation, surface normal prediction, point maps, and camera pose estimation, and supports language-defined task variants. Experiments show a single unified model rivals leading task-specific systems in structured visual understanding, dense geometry, segmentation, and multi-view geometry. Models and corpus are publicly released.

Paper Overview

Research Area: Computer Vision (CV) Authors: Xiaoyang Han, Jianhua Li, Kewang Deng Released: 2025-07-09 arXiv: 2507.06833

Key Idea

The authors formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model — with no task-specific architectures.

Under this formulation, SenseNova-Vision uses:

  • Natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions
  • Text outputs for symbolic predictions
  • Image outputs for dense spatial predictions
  • Mixed text-and-image outputs for compositional tasks
  • Training Data

    To support large-scale training, diverse computer vision annotations are converted into instruction-response examples compatible with these generation spaces, forming the SenseNova-Vision Corpus — a CV instruction-response corpus spanning text, image, and mixed targets.

    Training starts from an off-the-shelf pretrained unified multimodal model, using this corpus as the primary data source plus multimodal data as a capability-retention mix. No task-specific prediction heads or architecture modifications are required.

    Capabilities

    The resulting single model covers a wide range of visual tasks, including:

  • Detection
  • OCR
  • Keypoint estimation
  • Segmentation
  • Depth estimation
  • Surface normal prediction
  • Point maps
  • Camera pose estimation
It also supports language-defined task variants combining categories, colors, regions, and other visual cues.

Results

Experiments show that one unified model can match leading task-specific systems in structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry — suggesting unified multimodal generation is a scalable path to integrating computer vision into general-purpose foundation models.

The model and corpus are publicly released.

Original Abstract (excerpt)

> We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks. To support large-scale training, we convert diverse computer vision annotations into instruction-response examples compatible with these generation spaces, resulting in the SenseNova-Vision Corpus...

--- *Auto-collected on 2026-07-09*

Tags

#computer-vision#multimodal#unified-model#sensenova#arxiv#foundation-models#instruction-tuning#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346257