Summary
Canvas360 is a two-stage framework for in-context panoramic image generation that combines geometry-aware pretraining with task-specific fine-tuning, proposed by Haoran Feng, Ruiyang Zhang, and Longyi Zhang (arXiv 2507.08178). To overcome the scarcity of large-scale, high-quality training data for in-context panoramic tasks, the authors introduce Canvas360Dataset, a collection of 1 million paired panoramic samples covering style transfer, inpainting, outpainting, and editing, providing supervision across diverse generation scenarios. On the modeling side, Canvas360 strengthens text-to-panorama generation through parallel depth generation, velocity circular padding, and similarity loss regularization, which help the model learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence. Leveraging strong panoramic priors, Canvas360 offers a unified in-context framework that supports multiple downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility.
Overview
Field: Computer Vision
Authors: Haoran Feng, Ruiyang Zhang, Longyi Zhang
arXiv: 2507.08178
Abstract
In this work, the authors present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning.
To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, they propose Canvas360Dataset, a collection of 1M high-quality paired panoramic samples for style transfer, inpainting, outpainting, and editing, enabling effective supervision across diverse in-context generation scenarios.
On the modeling side, Canvas360 enhances text-to-panorama generation through:
- Parallel depth generation
- Velocity circular padding
- Similarity loss regularization
These enable the model to learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence.
Additionally, with strong panoramic priors, Canvas360 achieves a unified in-context panoramic generation framework that supports multiple downstream tasks via token-level concatenation, surpassing previous methods in both task coverage and modeling flexibility.
*Source: arXiv:2507.08178*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178346315