Summary
This paper by Lars van der Laan and Nathan Kallus (arXiv:2607.05375) introduces FORE (Fitted Occupancy-Ratio Evaluation), a fitted fixed-point method for estimating discounted occupancy ratios in off-policy evaluation for offline reinforcement learning. Occupancy ratios correct for distribution shift, and existing primal-dual and minimax approaches estimate them by enforcing occupancy-balance moments over critic classes. FORE instead characterizes the discounted occupancy ratio via an adjoint Bellman recursion: each iteration solves a single-layer density-ratio objective on one-step transition data, projecting the adjoint Bellman image onto a log-ratio class in KL divergence. Unlike fitted Q evaluation, whose analysis typically requires value-function realizability plus Bellman completeness or projector stability, FORE's core approximation condition is only realizability of the discounted occupancy ratio itself. The fitted ratio enables direct value estimation via reward reweighting, occupancy-weighted fitted Q evaluation, and doubly robust estimation combining the fitted ratio with a fitted Q function.
Paper Overview
- Field: Machine Learning
- Authors: Lars van der Laan, Nathan Kallus
- Published: 2026-07-06
- arXiv: 2607.05375
Summary
Occupancy ratios correct for distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over critic classes. This paper proposes Fitted Occupancy-Ratio Evaluation (FORE), a fitted fixed-point approach that characterizes the discounted occupancy ratio via an adjoint Bellman recursion.At each iteration, FORE solves a single-layer density-ratio objective on one-step transition data, projecting the adjoint Bellman image onto a log-ratio class in KL divergence.
Key Contribution
Whereas analyses of fitted Q evaluation typically require value-function realizability in addition to Bellman completeness or projector stability, FORE's core approximation condition is only realizability of the discounted occupancy ratio itself—no Bellman completeness is needed.The fitted ratio supports:
- Direct value estimation via reward reweighting
- Occupancy-weighted fitted Q evaluation
- Doubly robust estimation combining the fitted ratio with a fitted Q function
---
*Automatically collected on 2026-07-06*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178346224