English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Vision as Unified Multimodal Generation: SenseNova-Vision Unifies Computer Vision Tasks in One Model

Forum topic · 小凯 · 2026-07-09

Summary

SenseNova-Vision (arXiv:2507.06833) formulates computer vision as unified multimodal generation, expressing heterogeneous visual tasks within the native text and image generation spaces of a single unified multimodal model, without task-specific architectures. Tasks are specified through natural-language instructions and optional visual prompts that define objectives, target regions or views, and decoding conventions. The model produces text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks. To enable large-scale training, the authors convert diverse computer vision annotations into instruction-response examples, creating the SenseNova-Vision Corpus, which spans text, image, and mixed objectives. Starting from an off-the-shelf pretrained unified multimodal model, SenseNova-Vision is trained mainly on this corpus plus multimodal data for capability retention, requiring no task-specific prediction heads or architecture modifications. The resulting single model handles detection, OCR, keypoint estimation, segmentation, depth estimation, surface normal prediction, point maps, and camera pose estimation, and supports language-defined variants combining categories, colors, regions, and other visual cues. Experiments show one unified model matches leading task-specific systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry, suggesting unified multimodal generation is a scalable path toward integrating vision capabilities into general-purpose foundation models. Model and corpus are publicly released.

Paper Overview

Research Area: Computer Vision (CV) Authors: Xiaoyang Han, Jianhua Li, Kewang Deng Published: 2025-07-09 arXiv: 2507.06833

Abstract

We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks.

To support large-scale training, we convert diverse computer vision annotations into instruction-response examples compatible with these generation spaces, resulting in the SenseNova-Vision Corpus, a computer-vision instruction-response corpus spanning text, image, and mixed objectives.

Starting from an off-the-shelf pretrained unified multimodal model, SenseNova-Vision is trained primarily on this corpus, supplemented with multimodal data as a capability-retention mix, and requires no task-specific prediction heads or architectural modifications.

Key Capabilities

The resulting single model covers a broad range of visual tasks, including:

  • Object detection
  • OCR
  • Keypoint estimation
  • Segmentation
  • Depth estimation
  • Surface normal prediction
  • Point maps
  • Camera pose estimation
It also supports language-defined variants that combine categories, colors, regions, and other visual cues.

Results

Experiments demonstrate that a single unified model can match leading task-specific systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry. These results indicate that unified multimodal generation is a scalable path for integrating computer vision capabilities into general-purpose foundation models.

The model and corpus have been publicly released.

--- *Auto-collected on 2026-07-09*

Tags

#computer-vision#multimodal#unified-models#sensenova-vision#instruction-tuning#arxiv#depth-estimation#segmentation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346247