English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Directional Confusions Reveal Divergent Inductive Biases Through Rate-Distortion Geometry: Humans vs. Deep Vision Models

Forum topic · 小凯 · 2026-04-27

Summary

This paper, presented on zhichai.net, examines how humans and modern deep vision models make systematically different classification errors despite similar accuracy on natural image recognition. The authors—Leyla Roksan Caglar, Pedro A. M. Mediano, and Baihan Lin (arXiv:2604.21909)—analyze the directionality of confusions (which class is mistaken for which) across 12 perturbation types, using rate-distortion (RD) theory to characterize generalization geometry via slope (beta), curvature (kappa), and efficiency (AUC). Key findings: humans show broad but weak confusion asymmetries, while deep models exhibit sparser, stronger directional collapses; robustness training reduces global asymmetry but fails to restore human-like breadth-strength distributions; mechanistic simulations show different asymmetry organizations shift the RD frontier in opposite directions even at matched performance. The work positions directional confusion analysis and RD geometry as compact, interpretable signatures of inductive biases under distribution shift.

Paper Overview

  • Field: Computer Vision
  • Authors: Leyla Roksan Caglar, Pedro A. M. Mediano, Baihan Lin
  • Published: 2026-04-23
  • arXiv: 2604.21909
  • Abstract (English Translation)

    Humans and modern vision models can reach similar classification accuracy, yet they make systematically different types of errors—the difference lies not in error frequency, but in *who* is mistaken for *whom*, and in which direction. We show that these directional confusions reveal divergent inductive biases that accuracy alone cannot see.

    Using matched human and deep vision model responses in natural image classification across 12 perturbation types, we quantify asymmetries in confusion matrices and connect them, through the rate-distortion (RD) framework, to the geometry of generalization, summarized in three geometric features: slope (beta), curvature (kappa), and efficiency (AUC).

    We find that humans exhibit broad but weak asymmetries, whereas deep vision models show sparser, stronger directional collapses. Robustness training reduces global asymmetry but cannot recover the human-like breadth-strength distribution of graded similarity. Mechanistic simulations further show that different asymmetry organizations move the RD frontier in opposite directions, even when performance is matched. These results position directional confusions and RD geometry as compact, interpretable signatures of inductive biases under distribution shift.

    Key Takeaways

  • Error direction matters more than error rate when comparing human and machine vision.
  • Human confusions are broad but weak; deep model confusions are sparse and strong.
  • Robustness training reduces global asymmetry but does not fully humanize model error structure.
  • RD geometric features (slope, curvature, efficiency) shift in opposite directions depending on asymmetry organization—even at matched accuracy.
--- *Auto-collected on 2026-04-27*

Tags

#computer-vision#deep-learning#rate-distortion#inductive-bias#distribution-shift#confusion-matrix#arxiv-paper#human-vs-machine

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618801