Paper Overview
Field: Machine Learning Author: Javier Aguilar Martín Published: 2026-08-28 arXiv: 2608.28541
English Summary
A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. The paper characterizes what a certified model can know, and what its errors can cost, when the omission is an annular freeze mode enclosing an unreachable interior.
The gate quotient makes the question precise: acceptance-with-certainty determines the model exactly on the reachable query set; beyond reach is gauge. On a minimal ring instrument, the author proves the extreme case—a wrong-topology filled-disc artifact unfalsifiable by any sampling gate and bitwise harmless at play—and measures, with LLM synthesis across three model families, how one knob (a channel of width gamma) walks the same artifact through three regimes: unfalsifiable-and-harmless, falsifiable-and-costly, and immediately falsified.
Key Findings
1. Danger is topology relative to reach: A channel a planner can use collapses the blind model's exploitation (runtime cost drops from 1.09 to ~0, with a knee at gamma ~ 0.1), while a hidden channel with the same first Betti number keeps its full strength (1.12).
2. Repair is limited by parameters and sensors: No model family can recover the region from external evidence. From the interior, models propose the correct topology but cannot pin its parameters, and the proposed topology tracks the wrong beta_1 of the guided persistent-homology summary (a sensor with measured geometric resolution limits) rather than the truth.
3. Mitigation must match the error's dimension and direction: Point fences fail against one-dimensional boundaries; dimension-matched persistent fences collapse exploitation into a two-lesson transient (0.999 to 0.058), and dual freedom certificates symmetrically collapse invented-mode failures (1.769 to 0.029). In n dimensions, shells make misidentification nearly certain while danger remains fully exploitable: the two axes are independent.
Original Abstract (excerpt)
> A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. We characterize what a certified model can know, and what its errors can cost, when the omission is an annular freeze mode enclosing an unreachable interior. The gate quotient makes the question precise: acceptance-with-certainty determines the model exactly on the reachable query set; beyond reach is gauge. On a minimal ring instrument we prove the extreme case (a wrong-topology filled-disc artifact unfalsifiable by any sampling gate and bitwise harmless at play) and measure, with LLM synthesis across three model families, how one knob (a channel of width gamma) walks the same artifact through three regimes: unfalsifiable-and-harmless, falsifiable-and-costly,...
--- *Auto-collected on 2026-09-01*