Summary
This arXiv paper (2609.19135) by Pranaya Jajoo shows that off-policy evaluation (OPE) in POMDPs becomes exponentially hard when the logging policy depends on history. For every horizon H ≥ 3, the author constructs two POMDPs with at most two latent states per stage, three actions, and a common logger with three memory states. Despite satisfying action coverage, belief coverage, and behavior-marginal outcome-revealing conditions with constants independent of H, evaluating a known deterministic target policy to accuracy 1/8 at confidence 1-δ (δ ≤ 1/4) requires Θ((3/2)^H · log(1/δ)) logged episodes, even when both candidate models are fully known. The core mechanism is that a reset erases the unknown transition that determines the target policy's value. The author characterizes the induced statistical testing problem, derives matching optimal estimators, and validates the construction in a directed two-lane gridworld whose trajectory simulations match finite-sample predictions. The result establishes the intractability of history-dependent, model-based OPE under the behavior-marginal outcome-revealing definition of Zhang and Jiang (2025).
Paper Overview
Research Area: Machine Learning
Author: Pranaya Jajoo
Published: 2026-09-16
arXiv: 2609.19135
Key Contributions
Can a logged dataset visit every hidden state frequently and still be exponentially uninformative about a target policy's value? This paper shows that it can when the logger depends on history.
- Construction: For every horizon H ≥ 3, two POMDPs are constructed with at most two latent states per stage, three actions, and a common logger with three memory states.
- Coverage conditions hold: Action coverage, belief coverage, and two behavior-marginal outcome-revealing conditions all have constants independent of H.
- Exponential lower bound: Evaluating a known deterministic target policy to accuracy 1/8 requires Θ((3/2)^H · log(1/δ)) logged episodes at confidence 1−δ, for 0 < δ ≤ 1/4, even when both candidate models are known.
- Mechanism: A reset erases the unknown transition that determines the target value. The paper characterizes the resulting statistical testing problem exactly and derives matching optimal estimators.
- Empirical validation: A directed two-lane gridworld implements the construction, and trajectory simulations match the finite-sample predictions.
The result establishes intractability for the history-dependent, model-based setting under the behavior-marginal outcome-revealing definition of Zhang and Jiang (2025).
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178634945