English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

See & Sniff: Learning Visuo-Olfactory Representations with SmellNet-V

Forum topic · 小凯 · 2026-06-28

Summary

Researchers introduce SmellNet-V, a scalable visuo-olfactory dataset, and See & Sniff, a self-supervised framework for learning joint visual-olfactory representations. Since modern multimodal models rarely incorporate smell due to the lack of paired visuo-olfactory data, the authors exploit the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows synthetically pairing smell-only samples with semantically aligned in-the-wild web images, converting a unimodal olfactory dataset into a cross-modal benchmark without costly co-collection. The framework uses dense local alignment and naturally produces smell saliency maps for spatially grounding odor sources. The authors also introduce a pixel-level smell localization task and evaluation benchmark. Their approach outperforms olfaction-only baselines by 7% on olfactory classification and generalizes to cross-modal retrieval and smell localization.

Overview

  • Field: Computer Vision
  • Authors: Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Son Chung, Arda Senocak
  • Published: 2026-06-25
  • arXiv: 2606.27307
  • Abstract (translation)

    While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. The authors introduce SmellNet-V, a scalable visuo-olfactory dataset built on the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows synthetically pairing smell-only samples with semantically aligned in-the-wild web images, converting a unimodal olfactory dataset into a cross-modal benchmark without costly co-collection.

    Building on this dataset, the authors propose See & Sniff, a self-supervised framework that learns joint visuo-olfactory representations via dense local alignment and naturally produces smell saliency maps for spatial grounding of odor sources. They further introduce a pixel-level smell localization task and evaluation benchmark. Their method surpasses olfaction-only baselines by 7% on olfactory classification and generalizes to cross-modal retrieval and smell localization, establishing visuo-olfactory learning as a new direction in multimodal perception.

    Key contributions

  • SmellNet-V: a scalable cross-modal benchmark created by synthetically pairing olfactory data with web images
  • See & Sniff: self-supervised dense local alignment for joint visuo-olfactory representations
  • Smell saliency maps enabling spatial localization of odor sources
  • A new pixel-level smell localization task and benchmark
---

*Auto-collected on 2026-06-28*

Tags

#multimodal-learning#computer-vision#self-supervised-learning#olfaction#datasets#arxiv#smell-localization

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208240