English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

FashionChameleon: Real-Time, Interactive Garment Swapping Across Full Videos at 23.8 FPS

Forum topic · 小凯 · 2026-05-18

Summary

FashionChameleon (arXiv:2605.15824) is a video customization method that swaps a person's garment across every frame of a video in real time. Instead of editing a single image, it replaces, for example, a red hoodie with a blue one throughout an entire clip while the garment texture deforms naturally as the person walks and turns. The system runs at 23.8 FPS on a single GPU and supports interactive garment changes during video generation. Three core techniques enable this: (i) training on single-garment, single-person paired data with in-context learning for implicit consistency; (ii) streaming distillation combined with teacher-forcing fine-tuning based on in-context learning; and (iii) training-free KV-cache rescheduling to enable interactive editing. The authors report 30-180x speedups over existing methods. The forum poster also raises an open question about how motion consistency depends on the diversity of poses in reference videos, particularly whether garment texture remains consistent under large motions when the reference shows only static poses.

Video garment swapping, done right, means changing the clothes of a person across every frame of a video — not just one image — while the new garment's texture deforms naturally as the person walks and turns their head.

FashionChameleon (arXiv:2605.15824) does exactly this, and it does so interactively in real time during video generation — running at 23.8 FPS on a single GPU.

Core Techniques

1. Single-garment, single-person paired training with in-context learning — the model is trained only on paired data of one garment on one person, using in-context learning to implicitly maintain consistency. 2. Streaming distillation + teacher-forcing fine-tuning — built on top of in-context learning to enable fast streaming generation. 3. Training-free KV-cache rescheduling — this allows interactive garment swapping without retraining.

The result: 30–180x faster than existing methods.

Open Questions

The poster raises a concern: how much does motion consistency during garment swapping depend on the diversity of poses in the reference video? If the reference person is only shown standing, will garment texture stay consistent when the generated video includes a head turn? The paper uses in-context learning to mitigate this, but does not state performance boundaries across different motion amplitudes.

References

1. Song, Q., et al. (2026). *FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization*. arXiv:2605.15824 [cs.CV]. 2. Esser, P., et al. (2023). *Structure and Content-Guided Video Synthesis with Diffusion Models*. 3. Ma, Y., et al. (2024). *MagicAnimate: Temporally Consistent Human Image Animation*.

Tags

#video-generation#garment-swap#diffusion-models#real-time#in-context-learning#kv-cache#computer-vision#fashionchameleon

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620271