← All research notes

Quantum reinforcement learning / ICML 2025

Can quantum information help an agent learn better?

Q-UCRL: a theoretical route to less reward missed while learning.

The takeaway

Q-UCRL shows how access to quantum information about an environment can improve the theoretical learning guarantee for an agent making decisions over a long horizon.

01 / Learning without a finish line

Some decisions keep coming: there is no final level or natural reset. An agent must explore unfamiliar actions while earning rewards along the way. Regret measures the reward it misses compared with an optimal policy.

An RL agent receives ordinary outcomes and quantum information about possible next states
Two sources of information. The lower loop is ordinary interaction: take an action, observe an outcome. The upper loop supplies quantum information to estimate what could happen next. Figure 1, page 4.

02 / Turn better estimates into better decisions

The agent has access to a quantum transition oracle: a special interface encoding possible next states. Q-UCRL uses quantum mean estimation to learn these transition probabilities, then chooses an optimistic policy.

  1. 01Estimate

    Use quantum information to refine the model of the environment.

  2. 02Remember

    Combine fresh estimates with earlier estimates using visit counts.

  3. 03Explore

    Choose the best policy among plausible models, then repeat.

Measurement changes quantum states. The estimator carries forward the information already learned across training phases, rather than assuming measured states can simply be reused.

03 / What changes in the guarantee?

Classical reference√T

Square-root growth in the training horizon.

Q-UCRL with quantum accesspolylog(T)

Growth through powers of logarithms.

A schematic comparison of theoretical scaling, not an experimental plot. T is the number of interaction rounds; other problem factors and logarithmic terms in the classical reference are omitted.

A stronger information interface.

The gain comes from the assumed quantum access and a new analysis. It does not contradict classical lower bounds, which assume ordinary observations.

Why this work matters

It connects a quantum estimation advantage to an end-to-end reinforcement learning guarantee. The contribution includes the learning algorithm, an estimator that accounts for measurement, and a regret proof that avoids martingale concentration arguments.

Scope: This is a theoretical result for finite-state, finite-action average-reward problems under the paper’s quantum-oracle and mixing assumptions. It is not a hardware benchmark or a demonstrated wall-clock speedup for today’s RL systems.

Technical context & sources

The paper writes its horizon dependence as Õ(1), with logarithmic factors suppressed. This does not mean constant or zero regret: Theorem 1 includes powers of logarithms in T, as well as state-space, action-space, and mixing-time factors. Sections 3–5 define the access model, Q-UCRL, and the analysis; equations (21)–(25) describe the estimator’s count-weighted update.

The diagram is an unmodified crop of Figure 1 from the supplied paper. The learning-cycle and scaling panels are explanatory summaries created for this note.