Reference health

Importance sampling in reinforcement learning with an estimated behavior policy

https://doi.org/10.1007/s10994-020-05938-9
CiteStamped reference-health badge
24/24 checkable references clean · checked 2026-09-03

Every reference with a DOI in the deposited reference list resolved to a known work in Crossref or DataCite at the dated check, and none carried a retraction, withdrawal, or removal notice.

42 without a DOI — not checked. A reference deposited without a DOI is never matched by title or guessed at; it stays outside the checked set, and this line discloses that.

The 24 checked references that resolve
resolves10.1080/00273171.2011.568786
An Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies
resolves10.1126/science.153.3731.34
Dynamic Programming
resolves10.1609/aaai.v31i1.10810
OFFER: Off-Environment Reinforcement Learning
resolves10.1609/aaai.v32i1.11607
Expected Policy Gradients
resolves10.3150/15-BEJ725
Integral approximation by kernel smoothing
resolves10.24963/ijcai.2018/729
Importance Sampling for Fair Policy Selection
resolves10.1214/aoms/1177728174
Asymptotic Minimax Character of the Sample Distribution Function and of the Classical Multinomial Estimator
resolves10.1145/1390156.1390199
Reinforcement learning in the presence of rare events
resolves10.1609/aaai.v33i01.33013647
Off-Policy Deep Reinforcement Learning by Bootstrapping the Covariate Shift
resolves10.1007/978-94-009-5819-7
Monte Carlo Methods
resolves10.1093/biomet/asm076
Importance Sampling Via the Estimated Sampler
resolves10.1111/1468-0262.00442
Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score
resolves10.1017/CBO9780511809071
Introduction to Information Retrieval
resolves10.1038/nature14236
Human-level control through deep reinforcement learning
resolves10.2139/ssrn.3300346
Efficient Counterfactual Learning from Bandit Feedback
resolves10.1111/rssb.12185
Control Functionals for Monte Carlo Integration
resolves10.1016/j.neunet.2008.02.003
Reinforcement learning of motor skills with policy gradients
resolves10.1080/01621459.1987.10478441
Model-Based Direct Adjustment
resolves10.1016/B978-1-55860-307-3.50045-9
A Reinforcement Learning Method for Maximizing Undiscounted Rewards
resolves10.1038/nature16961
Mastering the game of Go with deep neural networks and tree search
resolves10.1609/aaai.v31i1.10932
Importance Sampling with Unequal Support
resolves10.1109/ADPRL.2009.4927542
A theoretical and empirical analysis of Expected Sarsa
resolves10.24963/ijcai.2018/414
A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning
resolves10.1007/978-3-030-63823-8
Neural Information Processing
The 42 references without a DOI — listed, not checked
no DOI — not checkedAsadi, K., Allen, C., Roderick, M., Mohamed, A.-R., Konidaris, G., & Littman, M. (2017). Mean actor critic. arXiv preprint arXiv:1709.00503v1.
no DOI — not checkedAsis, K. D., Hernandez-Garcia, J. F., Holland, G. Z., & Sutton, R. S. (2018). Multi-step reinforcement learning: A unifying algorithm. In Proceedings of the 32nd AAAI conference on artificial intelligence (AAAI).
no DOI — not checkedBrockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). OpenAI gym. arXiv preprint arXiv:1606.01540.
no DOI — not checkedDhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., & Wu, Y. (2017). OpenAI baselines. https://github.com/openai/baselines.
no DOI — not checkedDudík, M., Langford, J., & Li, L. (2011). Doubly robust policy evaluation and learning. In Proceedings of the 28th international conference on machine learning (ICML).
no DOI — not checkedFarajtabar, M., Chow, Y., & Ghavamzadeh, M. (2018). More robust doubly robust off-policy evaluation. In Proceedings of the 35th international conference on machine learning (ICML).
no DOI — not checkedFellows, M., Ciosek, K., & Whiteson, S. (2018). Fourier policy gradients. arXiv preprint arXiv:1802.06891.
no DOI — not checkedGreensmith, E., Bartlett, P. L., & Baxter, J. (2004). Variance reduction techniques for gradient estimates in reinforcement learning. Journal of Machine Learning Research, 5, 1471–1530.
no DOI — not checkedHallak, A., & Mannor, S. (2017). Consistent on-line off-policy evaluation. In Proceedings of the 34th international conference on machine learning, pp. 1372–1383.
no DOI — not checkedHanna, J. P., & Stone, P. (2019). Reducing sampling error in the monte carlo policy gradient estimator. In Proceedings of the 19th international conference on autonomous agents and multi-agent systems (AAMAS).
no DOI — not checkedHanna, J. P., Thomas, P. S., Stone, P., & Niekum, S. (2017). Data-efficient policy evaluation through behavior policy search. In Proceedings of the 34th international conference on machine learning (ICML).
no DOI — not checkedHanna, J. P., Niekum, S., & Stone, P. (2019). Importance sampling with an estimated behavior policy. In Proceedings of the 36th international conference on machine learning (ICML).
no DOI — not checkedJiang, N., & Li, L. (2016). Doubly robust off-policy evaluation for reinforcement learning. In Proceedings of the 33rd international conference on machine learning (ICML).
no DOI — not checkedKingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. In Proceedings of the international conference on learning representations (ICLR).
no DOI — not checkedLi, L., Munos, R., & Szepesvári, C. (2015). Toward minimax off-policy value estimation. In Proceedings of the 18th international conference on artificial intelligence and statistics.
no DOI — not checkedLiu, Q., & Lee, J. D. (2017). Black-box importance sampling. In Proceedings of the 20th international conference on artificial intelligence and statistics.
no DOI — not checkedLiu, Q., Li, L., Tang, Z., & Zhou, D. (2018). Breaking the curse of horizon: Infinite-horizon off-policy estimation. Advances in Neural Information Processing Systems (NeurIPS), 31, 5356–5366.
no DOI — not checkedMahadevan, S. (1996). Average reward reinforcement learning: Foundations, algorithms, and empirical results. Machine Learning, 1, 159–196.
no DOI — not checkedMnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd international conference on machine learning (ICML), pp. 1928–1937.
no DOI — not checkedMoore, A (1990). Efficient Memory-based learning for robot control. PhD thesis, University of Cambridge.
no DOI — not checkedMousavi, A., Li, L., Liu, Q., & Zhou, D. (2020). Black-box off-policy estimation for infinite-horizon reinforcement learning. In International conference on learning representations (ICLR)
no DOI — not checkedPavse, B. S., Durugkar, I., Hanna, J. P., & Stone, P. (2020). Reducing sampling error in batch temporal difference learning. In Proceedings of the 37th international conference on machine learning (ICML).
no DOI — not checkedPetit, B., Amdahl-Culleton, L., Liu, Y., Smith, J., & Bacon, P. L. (2019). All-action policy gradient methods: A numerical integration approach. arXiv preprint arXiv:1910.09093.
no DOI — not checkedPrecup, D., Sutton, R. S., & Singh, S. (2000). Eligibility traces for off-policy policy evaluation. In Proceedings of the 17th international conference on machine learning (ICML), pp. 759–766.
no DOI — not checkedPuterman, M. L. (2014). Markov decision processes: Discrete stochastic dynamic programming. New York: Wiley.
no DOI — not checkedRaghu, A., Gottesman, O., Liu, Y., Komorowski, M., Faisal, A., Doshi-Velez, F., & Brunskill, F. (2018). Behaviour policy estimation in off-policy policy evaluation: Calibration matters. In Proceedings of the ICML workshop on causal inference, counterfactual prediction, and autonomous action.
no DOI — not checkedRubinstein, R. Y., & Kroese, D. P. (2013). The cross-entropy method: a unified approach to combinatorial optimization Monte Carlo simulation and machine learning. Berlin: Springer.
no DOI — not checkedRummery, G. A., & Niranjan, M. (1994). On-line Q-learning using connectionist systems, Vol. 37. Department of Engineering Cambridge, England: University of Cambridge.
no DOI — not checkedSchulman, J., Levine, S., Moritz, P., Jordan, M., & Abbeel, P. (2015). Trust region policy optimization. In Proceedings of the 32nd international conference on machine learning (ICML). URL http://jmlr.csail.mit.edu/proceedings/papers/v37/schulman15.html.
no DOI — not checkedSchulman, J., Moritz, P., Levine, S., Jordan, M. I., & Abbeel, P. (2016). High-dimensional continuous control using generalized advantage estimation. In Proceedings of the international conference on learning representations (ICLR).
no DOI — not checkedSchulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
no DOI — not checkedShi, L., Li, S., Cao, L., Yang, L., & Pan, G. (2019). TBQ ($$\sigma$$): Improving efficiency of trace utilization for off-policy reinforcement learning. In Proceedings of the 18th international conference on autonomous agents and multiagent systems (AAMAS), pp. 1025–1032.
no DOI — not checkedSingh, S. P., & Sutton, R. S. (1996). Reinforcement learning with replacing eligibility traces. Machine Learning, 22, 123–158.
no DOI — not checkedSutton, R. S. (1984). Temporal credit assignment in reinforcement learning. PhD thesis, University of Massachusetts, Amherst.
no DOI — not checkedSutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction. Cambridge: MIT Press.
no DOI — not checkedSutton, R. S., Singh, S., & McAllester, D. (2000). Comparing policy-gradient algorithms.
no DOI — not checkedThomas, P. S. (2015). Safe reinforcement learning. PhD thesis, University of Massachusetts Amherst.
no DOI — not checkedThomas, P. S., & Brunskill, E. (2016a). Data-efficient off-policy policy evaluation for reinforcement learning. In Proceedings of the 33rd international conference on machine learning (ICML).
no DOI — not checkedThomas, P. S., & Brunskill, E. (2016b). Magical policy search: Data efficient reinforcement learning with guarantees of global optimality. In European workshop on reinforcement learning.
no DOI — not checkedThomas, P. S., Theocharous, G., & Ghavamzadeh, M. (2015). High confidence policy improvement. In Proceedings of the 32nd international conference on machine learning (ICML).
no DOI — not checkedWilliams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3–4), 229–256.
no DOI — not checkedXie, Y., Liu, B., Liu, Q., Wang, Z., Zhou, Y., & Peng, J. (2018). Off-policy evaluation and learning from logged bandit feedback: Error reduction via surrogate policy. In Proceedings of the international conference on learning representations (ICLR).
What this badge says. CiteStamped means the CHECKABLE references of this work were clean at the dated check: each resolved to a known work in a public registry, and none carried a retraction notice at that time. It says nothing about the quality, findings, or importance of the work itself, and nothing about references deposited without a DOI.

checked 2026-09-03 — re-checked daily as this page is visited; titles and statuses come from Crossref and DataCite and are not part of the signed record

Embed this badge

Both snippets point at the live badge image and link back to this page. The badge re-renders from the daily check, so an embed never goes stale by more than a day of visits.

<a href="https://citestamp.com/citestamped/10.1007/s10994-020-05938-9"><img src="https://citestamp.com/citestamped/10.1007/s10994-020-05938-9/badge.svg" alt="CiteStamped reference-health badge" width="460" height="64"></a>
[![CiteStamped reference-health badge](https://citestamp.com/citestamped/10.1007/s10994-020-05938-9/badge.svg)](https://citestamp.com/citestamped/10.1007/s10994-020-05938-9)