Every reference with a DOI in the deposited reference list resolved to a known
work in Crossref or DataCite at the dated check, and none carried a retraction,
withdrawal, or removal notice.
The 25 checked references that resolve
resolves10.1109/jas.2018.7511249Feature-based aggregation and deep reinforcement learning: a survey and some new implementations
resolves10.1214/aoms/1177729330A Measure of Asymptotic Efficiency for Tests of a Hypothesis Based on the sum of Observations
resolves10.1109/cdc.2006.377527Clinical data based optimal STI strategies for HIV: a reinforcement learning approach
resolves10.1111/1468-0262.00442Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score
resolves10.1145/1935826.1935878Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
resolves10.1109/embc.2016.7591355Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach
resolves10.1007/11564096_32Neural Fitted Q Iteration – First Experiences with a Data Efficient Neural Reinforcement Learning Method
The 34 references without a DOI — listed, not checked
no DOI — not checkedref1
no DOI — not checkedref4
no DOI — not checkedref6
no DOI — not checkedInput generalization in delayed reinforcement learning: An algorithm and performance comparisons
no DOI — not checkedref11
no DOI — not checkedref13
no DOI — not checkedref14
no DOI — not checkedProvably efficient rl with rich observations via latent state decoding
no DOI — not checkedref16
no DOI — not checkedTree-based batch mode reinforcement learning
no DOI — not checkedMetrics for finite markov decision processes
no DOI — not checkedOff-policy deep reinforcement learning without exploration
no DOI — not checkedref22
no DOI — not checkedref24
no DOI — not checkedThe optimal sample complexity of pac learning
no DOI — not checkedDoubly robust off-policy value evaluation for reinforcement learning
no DOI — not checkedProvably efficient reinforcement learning with linear function approximation
no DOI — not checkedEfficiently breaking the curse of horizon: Double reinforcement learning in infinite-horizon processes
no DOI — not checkedDouble reinforcement learning for efficient off-policy evaluation in markov decision processes
no DOI — not checkedref35
no DOI — not checkedStabilizing off-policy q-learning via bootstrapping error reduction
no DOI — not checkedTowards a unified theory of state abstraction for mdps
no DOI — not checkedBreaking the curse of horizon: Infinite-horizon off-policy estimation
no DOI — not checkedRepresentation balancing mdps for off-policy policy evaluation
no DOI — not checkedOffline policy evaluation across representations with applications to educational games
no DOI — not checkedKinematic state abstraction and provably efficient rich-observation reinforcement learning
no DOI — not checkedref46
no DOI — not checkedref49
no DOI — not checkedSmdp homomorphisms: An algebraic approach to abstraction in semi markov decision processes
no DOI — not checkedref53
no DOI — not checkedref56
no DOI — not checkedref58
no DOI — not checkedref60
no DOI — not checkedSolar: Deep structured representations for model-based reinforcement learning
checked 2026-08-27 — re-checked daily as this page is visited;
titles and statuses come from Crossref and DataCite and are not part of the signed record
Both snippets point at the live badge image and link back to this page. The
badge re-renders from the daily check, so an embed never goes stale by more than a day of visits.