Quantum reinforcement learning / ICML 2025
Can quantum information help an agent learn better?
Q-UCRL: a theoretical route to less reward missed while learning.
Q-UCRL shows how access to quantum information about an environment can improve the theoretical learning guarantee for an agent making decisions over a long horizon.
01 / Learning without a finish line
Some decisions keep coming: there is no final level or natural reset. An agent must explore unfamiliar actions while earning rewards along the way. Regret measures the reward it misses compared with an optimal policy.

02 / Turn better estimates into better decisions
The agent has access to a quantum transition oracle: a special interface encoding possible next states. Q-UCRL uses quantum mean estimation to learn these transition probabilities, then chooses an optimistic policy.
- 01Estimate
Use quantum information to refine the model of the environment.
- 02Remember
Combine fresh estimates with earlier estimates using visit counts.
- 03Explore
Choose the best policy among plausible models, then repeat.
Measurement changes quantum states. The estimator carries forward the information already learned across training phases, rather than assuming measured states can simply be reused.
03 / What changes in the guarantee?
Square-root growth in the training horizon.
Growth through powers of logarithms.
A schematic comparison of theoretical scaling, not an experimental plot. T is the number of interaction rounds; other problem factors and logarithmic terms in the classical reference are omitted.
The gain comes from the assumed quantum access and a new analysis. It does not contradict classical lower bounds, which assume ordinary observations.
Why this work matters
It connects a quantum estimation advantage to an end-to-end reinforcement learning guarantee. The contribution includes the learning algorithm, an estimator that accounts for measurement, and a regret proof that avoids martingale concentration arguments.
Scope: This is a theoretical result for finite-state, finite-action average-reward problems under the paper’s quantum-oracle and mixing assumptions. It is not a hardware benchmark or a demonstrated wall-clock speedup for today’s RL systems.
Technical context & sources
The paper writes its horizon dependence as Õ(1), with logarithmic factors suppressed. This does not mean constant or zero regret: Theorem 1 includes powers of logarithms in T, as well as state-space, action-space, and mixing-time factors. Sections 3–5 define the access model, Q-UCRL, and the analysis; equations (21)–(25) describe the estimator’s count-weighted update.
The diagram is an unmodified crop of Figure 1 from the supplied paper. The learning-cycle and scaling panels are explanatory summaries created for this note.