Paper Overview
Field: NLP Authors: Bohan Liu, Wenqian Ye, Guangzhi Xiong arXiv: 2507.03233
Background
Models trained via Contrastive Language-Image Pretraining (CLIP) are the foundational vision encoders for most modern Large Vision Language Models (LVLMs). However, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appearing within images confounds visual representations, biasing them toward lexical meaning rather than true visual semantics. This robustness issue, known as a Typographic Attack (TA), poses a significant risk to safety-critical applications such as autonomous driving.
Method
To achieve interpretable and effective robustness against TA, the authors propose a novel, training-free mechanistic interpretability approach:
- Provides sampling-based interpretations of hidden state representations
- Quantitatively attributes semantic versus lexical focus to individual attention heads
- Uses probabilistic analysis and circuit mining to isolate specific Vision Transformer (ViT) components that disproportionately encode lexical information, identifying the mechanical root of typographic attacks
- Simple interventions applied directly to the identified circuits—such as selectively adjusting attention weights—require no additional training yet significantly improve classification robustness against TA
- These interventions outperform both supervised and training-free defense methods
- When applied to the vision encoders of multiple state-of-the-art LVLMs, the approach substantially improves visual question answering (VQA) accuracy under typographic attacks on RIO-Bench, confirming the effectiveness and generalization of the mechanistic approach
- arXiv: https://arxiv.org/abs/2507.03233