Overview
- Field: Computer Vision
- Authors: Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Son Chung, Arda Senocak
- Published: 2026-06-25
- arXiv: 2606.27307
- SmellNet-V: a scalable cross-modal benchmark created by synthetically pairing olfactory data with web images
- See & Sniff: self-supervised dense local alignment for joint visuo-olfactory representations
- Smell saliency maps enabling spatial localization of odor sources
- A new pixel-level smell localization task and benchmark
Abstract (translation)
While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. The authors introduce SmellNet-V, a scalable visuo-olfactory dataset built on the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows synthetically pairing smell-only samples with semantically aligned in-the-wild web images, converting a unimodal olfactory dataset into a cross-modal benchmark without costly co-collection.
Building on this dataset, the authors propose See & Sniff, a self-supervised framework that learns joint visuo-olfactory representations via dense local alignment and naturally produces smell saliency maps for spatial grounding of odor sources. They further introduce a pixel-level smell localization task and evaluation benchmark. Their method surpasses olfaction-only baselines by 7% on olfactory classification and generalizes to cross-modal retrieval and smell localization, establishing visuo-olfactory learning as a new direction in multimodal perception.
Key contributions
*Auto-collected on 2026-06-28*