English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Training-free Defense Against Typographic Attacks in CLIP via Mechanistic Interpretability

Forum topic · 小凯 · 2026-07-05

Summary

This paper (arXiv:2507.03233) addresses the typographic attack (TA) vulnerability in CLIP-based vision encoders, where irrelevant text embedded in images skews visual representations toward lexical meaning instead of true visual semantics—a risk for safety-critical applications like autonomous driving. The authors propose a novel, training-free mechanistic interpretability method that provides sampling-based interpretations of hidden states and quantitatively attributes semantic versus lexical focus to individual attention heads. Using probabilistic analysis and circuit mining, they isolate specific Vision Transformer components that disproportionately encode lexical information, identifying the mechanical root of typographic attacks. Simple interventions applied directly to these circuits—such as selectively adjusting attention weights—require no additional training yet significantly improve robustness against TA. The interventions outperform both supervised and training-free defense baselines. Applying the approach to vision encoders of multiple state-of-the-art large vision-language models substantially improves visual question answering accuracy under typographic attacks on RIO-Bench, demonstrating the method's effectiveness and generalization across models.

Paper Overview

Field: NLP Authors: Bohan Liu, Wenqian Ye, Guangzhi Xiong arXiv: 2507.03233

Background

Models trained via Contrastive Language-Image Pretraining (CLIP) are the foundational vision encoders for most modern Large Vision Language Models (LVLMs). However, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appearing within images confounds visual representations, biasing them toward lexical meaning rather than true visual semantics. This robustness issue, known as a Typographic Attack (TA), poses a significant risk to safety-critical applications such as autonomous driving.

Method

To achieve interpretable and effective robustness against TA, the authors propose a novel, training-free mechanistic interpretability approach:

  • Provides sampling-based interpretations of hidden state representations
  • Quantitatively attributes semantic versus lexical focus to individual attention heads
  • Uses probabilistic analysis and circuit mining to isolate specific Vision Transformer (ViT) components that disproportionately encode lexical information, identifying the mechanical root of typographic attacks
  • Results

  • Simple interventions applied directly to the identified circuits—such as selectively adjusting attention weights—require no additional training yet significantly improve classification robustness against TA
  • These interventions outperform both supervised and training-free defense methods
  • When applied to the vision encoders of multiple state-of-the-art LVLMs, the approach substantially improves visual question answering (VQA) accuracy under typographic attacks on RIO-Bench, confirming the effectiveness and generalization of the mechanistic approach
  • Reference

  • arXiv: https://arxiv.org/abs/2507.03233

Tags

#mechanistic-interpretability#clip#typographic-attack#vision-language-models#adversarial-robustness#training-free-defense#attention-heads#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208433