English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Canvas360: Geometric-aware Pretraining for In-context Panoramic Generation (arXiv 2507.08178)

Forum topic · 小凯 · 2026-07-11

Summary

Canvas360 is a two-stage framework for in-context panoramic image generation that combines geometry-aware pretraining with task-specific fine-tuning, proposed by Haoran Feng, Ruiyang Zhang, and Longyi Zhang (arXiv 2507.08178). To overcome the scarcity of large-scale, high-quality training data for in-context panoramic tasks, the authors introduce Canvas360Dataset, a collection of 1 million paired panoramic samples covering style transfer, inpainting, outpainting, and editing, providing supervision across diverse generation scenarios. On the modeling side, Canvas360 strengthens text-to-panorama generation through parallel depth generation, velocity circular padding, and similarity loss regularization, which help the model learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence. Leveraging strong panoramic priors, Canvas360 offers a unified in-context framework that supports multiple downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility.

Overview

Field: Computer Vision Authors: Haoran Feng, Ruiyang Zhang, Longyi Zhang arXiv: 2507.08178

Abstract

In this work, the authors present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning.

To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, they propose Canvas360Dataset, a collection of 1M high-quality paired panoramic samples for style transfer, inpainting, outpainting, and editing, enabling effective supervision across diverse in-context generation scenarios.

On the modeling side, Canvas360 enhances text-to-panorama generation through:

  • Parallel depth generation
  • Velocity circular padding
  • Similarity loss regularization
These enable the model to learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence.

Additionally, with strong panoramic priors, Canvas360 achieves a unified in-context panoramic generation framework that supports multiple downstream tasks via token-level concatenation, surpassing previous methods in both task coverage and modeling flexibility.

*Source: arXiv:2507.08178*

Tags

#canvas360#panoramic-generation#computer-vision#diffusion-models#inpainting#text-to-panorama#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346315