English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Routing Is Least Learnable Where It Is Most Valuable: Upper and Lower Bounds for Web Agent Observation Modes

Forum topic · ✨步子哥 · 2026-08-09

Summary

A detailed analysis of the arXiv paper 2608.06171, which studies observation-mode routing for Web Agents. The paper tests six observation modes (text, pixels, both, 2x2, text+2corners, pixels+2corners) across 8 site-model combinations on VisualWebArena and WebArena, finding that modes are complementary, that the optimal mode reverses across task sets, and that Oracle upper bounds are inflated by 12-14% run-to-run noise. Its central structural finding: routing supervision labels are only produced when the agent succeeds, so label supply correlates negatively (r=0.95) with routing value—routing is hardest to learn exactly where it matters most. No tested routing policy reliably beats a fixed mode. As a practical alternative, the paper proposes a cost floor: run each task first with the cheapest mode (text) and escalate on failure, saving 9.5-30.6% of cost without router training. The post evaluates the paper's limitations and its broader implications for evaluation, training, and engineering.

A Counterintuitive Finding

Imagine managing a customer service team where each agent can work in three modes:

1. Text transcript only (fast, cheap, but misses tone) 2. Video recording only (most complete, but slow and expensive) 3. Text + video (most comprehensive, but costliest)

Your task: for each call, choose a mode. Intuitively, you'd train a router by trying all modes on every call and learning which mode works where. But a paper from August 2026 shows this approach has a fatal structural problem: the labels a router needs are only generated when the agent succeeds. The weaker the agent, the fewer labels—and the weaker the agent, the more routing is needed.

Routing is hardest to learn exactly where it is most valuable.

This is the core finding of Xian Sun et al.'s arXiv paper 2608.06171, *Routing Is Least Learnable Where It Is Most Valuable*.

routing-least-learnable.svg

Six Observation Modes

The paper studies observation-mode selection for Web Agents, which decide actions (click, type, scroll) based on page state. Six modes are compared:

1. text: page text content (HTML-to-text) 2. pixels: page screenshot 3. both: text + screenshot 4. 2x2: text + screenshot + zoomed corner crops 5. text+2corners: text + corner crops 6. pixels+2corners: screenshot + corner crops

Each mode has trade-offs: text is fast but lacks visual detail, pixels is slow but complete. No single mode dominates on all tasks.

Three Key Findings

Finding 1: Modes Are Complementary

Testing six modes on 8 site-model combinations from VisualWebArena and WebArena, the paper shows every mode solves tasks that other modes miss. Modes don't replace each other; they cover each other's blind spots.

Finding 2: The Optimal Mode Reverses Across Task Sets

The best mode on task set A can be the worst on task set B. Across the 8 site-model combinations, the optimal mode choice reverses. There is no universally optimal mode—a fixed mode is suboptimal, so a router must identify task characteristics and select accordingly.

Finding 3: The Oracle Upper Bound Is Inflated by Noise

Web Agent runs are stochastic: 12-14% of outcomes change on reruns. This means reported Oracle upper bounds (the assumption that a router always picks the optimal mode) are inflated. Since many routing evaluations use the Oracle bound as a reference, prior evaluations may be unreliable.

The Structural Dilemma of Routing

The paper tests 5 routing strategies (LLM-based, feature-similarity-based, historical-success-rate-based). None reliably beats always picking one fixed mode.

The core contradiction:

  • Stronger agent → more successes → more routing labels → easier router training
  • Weaker agent → fewer successes → fewer labels → harder router training
  • Routing is most valuable where the agent is weakest, but labels are scarcest there. The author measured the correlation between label supply and routing value: r = 0.95, an almost perfect negative correlation—analogous to the "Matthew effect" in education, where high performers accumulate resources while those most in need get the least.

    The Cost Floor: A Practical Engineering Solution

    Despite the difficulty of router training, the paper offers a practical scheme: run each task first with the cheapest mode (text); if it fails, escalate to a more expensive mode.

    This reduces cost on all 8 site-model combinations: 9.5-30.6% cost savings. It's achievable and requires no router training—though it sacrifices latency in the worst case.

    Implications at Three Levels

    1. Evaluation: The 12-14% noise inflation means prior Oracle upper bounds were overstated. Future evaluations must report noise levels. 2. Training: The r=0.95 negative correlation imposes a structural ceiling on router training—a data problem, not an algorithm problem. 3. Engineering: The cost-floor scheme delivers 9.5-30.6% savings without training.

    Cross-Paper Resonance

    The findings align with several recent papers:

  • Regression Tax: skills can degrade agents; routing can too, when trained poorly—"tools meant to help can backfire."
  • Looping Is Not Reliability: correctness isn't an absorbing state (82.0%→67.3%); routing's 12-14% noise is isomorphic—"past success ≠ current success."
  • TriviaRoomQA: models collapse past knowledge boundaries; routing is only needed at those boundaries, where it also learns worst.
  • MIST: evaluation blind spots hide problems—fixed modes look "stable," but Oracle bounds are inflated.
The shared theme: measurement coverage matters more than measurement depth.

An Honest Assessment

Limitations: (1) only Web Agent scenarios are tested, though the label-supply contradiction is likely structural; (2) the 5 routing policies may not represent the true ceiling of routers; (3) the cost floor trades latency for savings.

These don't undermine the core contribution: a structural dilemma—routing label supply is negatively correlated (r=0.95) with routing value—is a data problem, not an algorithm problem.

A Deeper Insight

The most memorable takeaway: the label supply for supervised learning can be negatively correlated with its value. This appears in RL (sparse rewards where improvement is most needed), active learning (uncertain samples are costliest to label), and education (struggling students get the least feedback). "Hardest to learn where most needed" is a structural dilemma, not an engineering bug.

Future breakthroughs may come not from better routing algorithms but from smarter label generation—synthetic data, counterfactual reasoning, or cross-agent transfer learning.

Conclusion

This paper elevates an engineering problem (routers train poorly) into a structural discovery (label supply is anti-correlated with value). Routing is least learnable where it is most valuable—applicable to any setting where supervision signals are only produced on success. When a router, classifier, or recommender trains poorly, ask: is its label supply negatively correlated with its value? If so, the problem isn't the algorithm—it's the structure.

---

Paper: https://arxiv.org/abs/2608.06171 HTML full text: https://arxiv.org/html/2608.06171v1

Tags

#web-agents#routing#observation-modes#machine-learning#evaluation#oracle-bounds#cost-optimization#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178603088