English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AI Thinks "Her" but Says "He": Vision-Language Models Suppress Female Representations Under Ambiguous Input

Forum topic · 小凯 · 2026-06-01

Summary

A Harvard study (arXiv:2605.31556) by Arnau Marin-Llobet, Simon Henniger, and implicit-bias researcher Mahzarin R. Banaji reveals a systematic gender bias inside vision-language models (VLMs). Using a new probing method called LALS (Latent Association Leaning Score), the authors project visual-token activations into the model's text embedding space and track gender associations layer by layer across 800+ gender-ambiguous images spanning 15 occupations and 4 mainstream VLMs. The key finding: when gender is not visible in an image, models often encode female associations internally—signals that peak in the middle layers, sometimes exceeding male associations—but these signals are systematically suppressed before output generation, while male associations pass through amplified. The authors call this an "asymmetric filter." Notably, clothing colors modulate the effect, suggesting culturally loaded visual cues act as gendered signal carriers. Alignment training suppresses biased outputs when gender is visible but leaves internal representations untouched; under ambiguity, models revert to a male default. The paper maps where bias emerges but does not propose fixes.

AI Thinks "Her" but Says "He": VLMs Suppress Female Representations Under Ambiguous Input

| Item | Detail | |------|--------| | Paper | Vision-Language Models Suppress Female Representations Under Ambiguous Input | | Authors | Arnau Marin-Llobet, Simon Henniger, Mahzarin R. Banaji | | Institution | Harvard University (Banaji is a co-creator of the IAT and a foundational figure in implicit-bias research) | | arXiv ID | 2605.31556 | | Submitted | May 29, 2026 | | Categories | cs.CV + cs.AI + cs.CL + cs.CY + cs.HC | | Core finding | Under gender-ambiguous inputs (full protective gear, figures seen from behind), VLMs default to male outputs—even when internal representations encode female associations. Male signals are amplified from input to output; female signals peak in mid-network layers and are suppressed before generation. This "asymmetric filtering" appears consistently across 15 occupations, 800+ ambiguous images, and 4 VLMs. |

1. A Figure Seen from Behind

A person in full protective gear and a helmet stands at a construction site. No face, no visible body shape—just someone doing something. Ask a VLM to caption it and it writes: "A worker is inspecting the equipment"—in the masculine. Another image: someone bent over, reading to children in a classroom. Again, "The teacher is reading to the children"—masculine.

The person in the first image might be a female engineer; the second, a male preschool teacher. But across 800+ gender-ambiguous images, 15 occupations, and 4 mainstream VLMs, the paper finds a systematic asymmetry:

When the input is ambiguous, VLMs internally encode "female"—yet the output collapses into "male."

2. The LALS Probe

The paper introduces LALS—Latent Association Leaning Score. As visual tokens pass through Transformer layers, each layer has internal activations. LALS projects these visual-token activations into the model's own text embedding space, then measures how close each activation point sits to male versus female concepts.

It is effectively a mirror placed at every layer—not asking the model what it saw, but peeking at what it is "thinking." LALS tracked gender associations layer by layer, token by token, across four VLMs and 15 occupation-related ambiguous images.

3. Internal Signals and External Outputs Are Two Different Ledgers

The central finding: when gender is invisible in the input, VLM internal representations frequently encode female associations—especially for traditionally female-dominated occupations (nurse, preschool teacher, administrative assistant). LALS shows strong mid-layer signals of "female" proximity.

Yet the final descriptions almost invariably use "he." The model "thinks female" around layers 8–12; by layers 18–20, that signal is pressed down—and by the output layer, it has collapsed to male.

This is not the model failing to see. It is seeing—then washing it away.

4. The Asymmetric Filter

Layer-by-layer tracking yields a clear picture:

  • Male signals: pass through the entire network nearly undamped—amplified at every layer, flowing unobstructed from pixels to text.
  • Female signals: peak in the middle layers (~6–14), at times exceeding male signals—but are systematically attenuated after roughly layer 15, leaving only traces at the generation layer.
  • The paper calls this an "asymmetric filter." The shape is stable: the same "female peak + decay, male accumulation + amplification" pattern appears across 4 VLMs and 15 occupations.

    A chilling detail: clothing color modulates the process. In color-ablation experiments, removing color weakened internal female associations. Color itself is a culturally loaded signal carrier—pink and soft palettes, associated with women in training data, have become conditioned cues in the visual encoder.

    5. Alignment Solves Half the Problem

    The paper opens with: "Alignment teaches VLMs to avoid expressing demographic biases." When gender is clearly visible—faces and body shapes legible—today's models largely avoid overt gender bias in output. Alignment works, at least on the surface.

    But when gender is invisible, all that protection peels away and the model reverts to its default: male. This exposes the essence of alignment—it optimizes outputs, not internal representations. As long as the training objective never penalizes a mid-layer concept association, that association stays intact. Alignment polishes the mouth but doesn't touch the brain.

    The paper's finding parallels the concept-first paper (arXiv:2605.22007), which reported that 16–47% of hallucinations stem from "knowing but not choosing." Both reveal the same pattern: the model internally encodes the correct answer, but the training objective locks in a different output—there at the level of factual knowledge, here at the level of social bias.

    6. Open Questions

    The findings are profound, but some things remain unverified:

  • LALS's precision. Projecting visual tokens into text embedding space is a significant simplification; the projection may lose information or introduce spurious associations. "Zero-shot" typically trades away accuracy—whether the measured female signal reflects genuine internal concept activation or a known but vision-irrelevant co-occurrence pattern is unclear.
  • The origin of the asymmetric filter. The paper observes the phenomenon but offers no causal account. Is it bias in visual pretraining data, in language-alignment data, or an architectural inductive bias? Unanswered.
  • External validity. Do the 15 occupations balance male-dominated, female-dominated, and neutral categories? Do 800+ images cover enough scene diversity across countries, environments, and photographic angles? An excellent initial study—but short of a full audit of gender bias in VLMs.

7. The Limits of Alignment

Compressed to one sentence: alignment can teach a model not to say it—but not to stop thinking it.

The asymmetric filter is not written into code; it crystallizes from the joint forces of data, architecture, and training objectives. It is emergent bias. Fixing it requires changing not prompts, not RLHF, but the training data, objectives, and architecture behind those filters.

The paper offers no fix. But it offers something equally important: a precise map of the layer where the problem is manufactured. With the map, repair becomes possible.

References

1. Marin-Llobet, Henniger & Banaji, "Vision-Language Models Suppress Female Representations Under Ambiguous Input", arXiv:2605.31556, 2026. 2. Bai et al., "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback", arXiv:2204.05862, 2022. 3. Birhane et al., "Multimodal Datasets: Misogyny, Pornography, and Malignant Stereotypes", arXiv:2110.01963, 2021. 4. Wolfe et al., "Concept-First: Commitment Failures in Large Language Models", arXiv:2605.22007, 2026. 5. Greenwald & Banaji, "Implicit Social Cognition: Attitudes, Self-Esteem, and Stereotypes", Psychological Review, 1995.

Tags

#vision-language-models#gender-bias#asymmetric-filter#alignment#internal-representations#implicit-bias#model-probing#ai-ethics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980711