Every reference with a DOI in the deposited reference list resolved to a known
work in Crossref or DataCite at the dated check, and none carried a retraction,
withdrawal, or removal notice.
The 42 references without a DOI — listed, not checked
no DOI — not checkedAsadi, K., Allen, C., Roderick, M., Mohamed, A.-R., Konidaris, G., & Littman, M. (2017). Mean actor critic. arXiv preprint arXiv:1709.00503v1.
no DOI — not checkedAsis, K. D., Hernandez-Garcia, J. F., Holland, G. Z., & Sutton, R. S. (2018). Multi-step reinforcement learning: A unifying algorithm. In Proceedings of the 32nd AAAI conference on artificial intelligence (AAAI).
no DOI — not checkedBrockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). OpenAI gym. arXiv preprint arXiv:1606.01540.
no DOI — not checkedDhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., & Wu, Y. (2017). OpenAI baselines. https://github.com/openai/baselines.
no DOI — not checkedDudík, M., Langford, J., & Li, L. (2011). Doubly robust policy evaluation and learning. In Proceedings of the 28th international conference on machine learning (ICML).
no DOI — not checkedFarajtabar, M., Chow, Y., & Ghavamzadeh, M. (2018). More robust doubly robust off-policy evaluation. In Proceedings of the 35th international conference on machine learning (ICML).
no DOI — not checkedFellows, M., Ciosek, K., & Whiteson, S. (2018). Fourier policy gradients. arXiv preprint arXiv:1802.06891.
no DOI — not checkedGreensmith, E., Bartlett, P. L., & Baxter, J. (2004). Variance reduction techniques for gradient estimates in reinforcement learning. Journal of Machine Learning Research, 5, 1471–1530.
no DOI — not checkedHallak, A., & Mannor, S. (2017). Consistent on-line off-policy evaluation. In Proceedings of the 34th international conference on machine learning, pp. 1372–1383.
no DOI — not checkedHanna, J. P., & Stone, P. (2019). Reducing sampling error in the monte carlo policy gradient estimator. In Proceedings of the 19th international conference on autonomous agents and multi-agent systems (AAMAS).
no DOI — not checkedHanna, J. P., Thomas, P. S., Stone, P., & Niekum, S. (2017). Data-efficient policy evaluation through behavior policy search. In Proceedings of the 34th international conference on machine learning (ICML).
no DOI — not checkedHanna, J. P., Niekum, S., & Stone, P. (2019). Importance sampling with an estimated behavior policy. In Proceedings of the 36th international conference on machine learning (ICML).
no DOI — not checkedJiang, N., & Li, L. (2016). Doubly robust off-policy evaluation for reinforcement learning. In Proceedings of the 33rd international conference on machine learning (ICML).
no DOI — not checkedKingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. In Proceedings of the international conference on learning representations (ICLR).
no DOI — not checkedLi, L., Munos, R., & Szepesvári, C. (2015). Toward minimax off-policy value estimation. In Proceedings of the 18th international conference on artificial intelligence and statistics.
no DOI — not checkedLiu, Q., & Lee, J. D. (2017). Black-box importance sampling. In Proceedings of the 20th international conference on artificial intelligence and statistics.
no DOI — not checkedLiu, Q., Li, L., Tang, Z., & Zhou, D. (2018). Breaking the curse of horizon: Infinite-horizon off-policy estimation. Advances in Neural Information Processing Systems (NeurIPS), 31, 5356–5366.
no DOI — not checkedMahadevan, S. (1996). Average reward reinforcement learning: Foundations, algorithms, and empirical results. Machine Learning, 1, 159–196.
no DOI — not checkedMnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd international conference on machine learning (ICML), pp. 1928–1937.
no DOI — not checkedMoore, A (1990). Efficient Memory-based learning for robot control. PhD thesis, University of Cambridge.
no DOI — not checkedMousavi, A., Li, L., Liu, Q., & Zhou, D. (2020). Black-box off-policy estimation for infinite-horizon reinforcement learning. In International conference on learning representations (ICLR)
no DOI — not checkedPavse, B. S., Durugkar, I., Hanna, J. P., & Stone, P. (2020). Reducing sampling error in batch temporal difference learning. In Proceedings of the 37th international conference on machine learning (ICML).
no DOI — not checkedPetit, B., Amdahl-Culleton, L., Liu, Y., Smith, J., & Bacon, P. L. (2019). All-action policy gradient methods: A numerical integration approach. arXiv preprint arXiv:1910.09093.
no DOI — not checkedPrecup, D., Sutton, R. S., & Singh, S. (2000). Eligibility traces for off-policy policy evaluation. In Proceedings of the 17th international conference on machine learning (ICML), pp. 759–766.
no DOI — not checkedPuterman, M. L. (2014). Markov decision processes: Discrete stochastic dynamic programming. New York: Wiley.
no DOI — not checkedRaghu, A., Gottesman, O., Liu, Y., Komorowski, M., Faisal, A., Doshi-Velez, F., & Brunskill, F. (2018). Behaviour policy estimation in off-policy policy evaluation: Calibration matters. In Proceedings of the ICML workshop on causal inference, counterfactual prediction, and autonomous action.
no DOI — not checkedRubinstein, R. Y., & Kroese, D. P. (2013). The cross-entropy method: a unified approach to combinatorial optimization Monte Carlo simulation and machine learning. Berlin: Springer.
no DOI — not checkedRummery, G. A., & Niranjan, M. (1994). On-line Q-learning using connectionist systems, Vol. 37. Department of Engineering Cambridge, England: University of Cambridge.
no DOI — not checkedSchulman, J., Levine, S., Moritz, P., Jordan, M., & Abbeel, P. (2015). Trust region policy optimization. In Proceedings of the 32nd international conference on machine learning (ICML). URL http://jmlr.csail.mit.edu/proceedings/papers/v37/schulman15.html.
no DOI — not checkedSchulman, J., Moritz, P., Levine, S., Jordan, M. I., & Abbeel, P. (2016). High-dimensional continuous control using generalized advantage estimation. In Proceedings of the international conference on learning representations (ICLR).
no DOI — not checkedSchulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
no DOI — not checkedShi, L., Li, S., Cao, L., Yang, L., & Pan, G. (2019). TBQ ($$\sigma$$): Improving efficiency of trace utilization for off-policy reinforcement learning. In Proceedings of the 18th international conference on autonomous agents and multiagent systems (AAMAS), pp. 1025–1032.
no DOI — not checkedSingh, S. P., & Sutton, R. S. (1996). Reinforcement learning with replacing eligibility traces. Machine Learning, 22, 123–158.
no DOI — not checkedSutton, R. S. (1984). Temporal credit assignment in reinforcement learning. PhD thesis, University of Massachusetts, Amherst.
no DOI — not checkedSutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction. Cambridge: MIT Press.
no DOI — not checkedSutton, R. S., Singh, S., & McAllester, D. (2000). Comparing policy-gradient algorithms.
no DOI — not checkedThomas, P. S. (2015). Safe reinforcement learning. PhD thesis, University of Massachusetts Amherst.
no DOI — not checkedThomas, P. S., & Brunskill, E. (2016a). Data-efficient off-policy policy evaluation for reinforcement learning. In Proceedings of the 33rd international conference on machine learning (ICML).
no DOI — not checkedThomas, P. S., & Brunskill, E. (2016b). Magical policy search: Data efficient reinforcement learning with guarantees of global optimality. In European workshop on reinforcement learning.
no DOI — not checkedThomas, P. S., Theocharous, G., & Ghavamzadeh, M. (2015). High confidence policy improvement. In Proceedings of the 32nd international conference on machine learning (ICML).
no DOI — not checkedWilliams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3–4), 229–256.
no DOI — not checkedXie, Y., Liu, B., Liu, Q., Wang, Z., Zhou, Y., & Peng, J. (2018). Off-policy evaluation and learning from logged bandit feedback: Error reduction via surrogate policy. In Proceedings of the international conference on learning representations (ICLR).
checked 2026-09-03 — re-checked daily as this page is visited;
titles and statuses come from Crossref and DataCite and are not part of the signed record
Both snippets point at the live badge image and link back to this page. The
badge re-renders from the daily check, so an embed never goes stale by more than a day of visits.