English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Canvas360: Enhancing In-context Panoramic Generation via Geometric-aware Pretraining and Task-specific Fine-tuning

Forum topic · 小凯 · 2026-07-12

Summary

Canvas360 is a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning, presented by researchers including Haoran Feng and Lu Qi. To address the scarcity of large-scale, high-quality training data for in-context panoramic tasks, the authors introduce Canvas360Dataset, a collection of 1 million high-quality paired panoramic samples supporting style transfer, inpainting, outpainting, and editing. On the modeling side, Canvas360 improves text-to-panorama generation through parallel depth generation, velocity circular padding, and similarity loss regularization, enabling the model to learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence. Leveraging strong panoramic priors, Canvas360 provides a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing prior methods in task coverage and modeling flexibility. Experiments show improved panorama fidelity, with particularly strong results on the panorama-specific FAED metric, achieving competitive or state-of-the-art results across reported evaluations.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Haoran Feng, Ruiyang Zhang, Longyi Zhang, Dizhe Zhang, Lu Qi
  • arXiv: 2607.08765
  • Project page: https://zry000.github.io/Canvas360/
  • Abstract

    In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning. To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, we propose Canvas360Dataset, a collection of 1M high-quality paired panoramic samples for style transfer, inpainting, outpainting, and editing, enabling effective supervision across diverse in-context generation scenarios. On the modeling side, Canvas360 enhances text-to-panorama generation through parallel depth generation, velocity circular padding, and similarity loss regularization, enabling the model to learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence. Furthermore, with strong panoramic priors, Canvas360 achieves a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing previous methods in task coverage and modeling flexibility. Extensive experiments demonstrate that Canvas360 improves panoramic image fidelity, with particularly notable performance on the panorama-specific FAED metric, achieving competitive or leading results across all reported quantitative evaluations.

    Key Contributions

  • Canvas360Dataset: 1M high-quality paired panoramic samples covering style transfer, inpainting, outpainting, and editing tasks.
  • Geometry-aware modeling: parallel depth generation, velocity circular padding, and similarity loss regularization for better geometric consistency.
  • Unified in-context framework: diverse downstream tasks enabled through token-level concatenation.
*Auto-collected on 2026-07-12.*

Tags

#canvas360#panoramic-generation#computer-vision#text-to-panorama#inpainting#outpainting#geometry-aware#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379391