Paper Overview
- Field: Computer Vision
- Authors: Leyla Roksan Caglar, Pedro A. M. Mediano, Baihan Lin
- Published: 2026-04-23
- arXiv: 2604.21909
- Error direction matters more than error rate when comparing human and machine vision.
- Human confusions are broad but weak; deep model confusions are sparse and strong.
- Robustness training reduces global asymmetry but does not fully humanize model error structure.
- RD geometric features (slope, curvature, efficiency) shift in opposite directions depending on asymmetry organization—even at matched accuracy.
Abstract (English Translation)
Humans and modern vision models can reach similar classification accuracy, yet they make systematically different types of errors—the difference lies not in error frequency, but in *who* is mistaken for *whom*, and in which direction. We show that these directional confusions reveal divergent inductive biases that accuracy alone cannot see.
Using matched human and deep vision model responses in natural image classification across 12 perturbation types, we quantify asymmetries in confusion matrices and connect them, through the rate-distortion (RD) framework, to the geometry of generalization, summarized in three geometric features: slope (beta), curvature (kappa), and efficiency (AUC).
We find that humans exhibit broad but weak asymmetries, whereas deep vision models show sparser, stronger directional collapses. Robustness training reduces global asymmetry but cannot recover the human-like breadth-strength distribution of graded similarity. Mechanistic simulations further show that different asymmetry organizations move the RD frontier in opposite directions, even when performance is matched. These results position directional confusions and RD geometry as compact, interpretable signatures of inductive biases under distribution shift.