← All research notes

Parallel reinforcement learning / Communication efficiency

Can agents learn together without constantly talking?

DIST-UCRL: share experience when it is time to update, rather than after every action.

The takeaway

Parallel agents can benefit from each other’s experience without constantly synchronizing. DIST-UCRL uses a count-based trigger for sharing, with learning guarantees and substantially fewer communication rounds in the reported experiments.

01 / The problem

Imagine several agents practicing the same task independently. Sharing experience helps them learn together, but checking in after every action creates communication overhead. How much coordination is actually necessary?

Parallel workers share interaction statistics through a coordinator and receive a common policy.
Many workers, shared learning. A conceptual schematic from the research profile: each worker interacts with its own independent copy of the same environment.

02 / The idea

Make communication depend on collected experience. Each agent tracks how often it tries an action in a particular situation. When a count reaches a threshold tied to the previously shared total, it requests a group synchronization. A coordinator pools statistics, updates the policy, and sends it back.

The agents then continue independently until the next trigger. The novelty is the synchronization rule and its analysis—not simply running more agents.

03 / Does it help?

We evaluated RiverSwim, an extended RiverSwim task, and Gridworld, averaging 50 independent runs. Below are two examples pairing learning performance with communication cost.

Two questions, side by side

Learning: lower regret means less reward missed compared with optimal behavior. Communication: lower counts mean fewer group check-ins. M is the number of agents; the legends differ between the two plot types.

RiverSwim per-agent regret comparison
RiverSwim · Learning. DIST-UCRL closely tracks the frequent-communication baseline for the same number of agents. Figure 1(a).
RiverSwim synchronization rounds
RiverSwim · Communication. Check-ins increase slowly as training continues. All curves here are DIST-UCRL. Figure 2(a).
Gridworld per-agent regret comparison
Gridworld · Learning. The similar curves show the learning benefit is retained in this experiment. Figure 1(c).
Gridworld synchronization rounds
Gridworld · Communication. Hundreds of synchronization rounds over one million time steps, rather than a check-in every step. Figure 2(c).

Select any figure to open its full-size view.

Similar learning. Fewer check-ins.

DIST-UCRL matches the regret-bound scaling of the paper’s frequent-communication comparator, while synchronization grows only logarithmically with training time for a fixed problem size.

Why this work matters

It treats communication as a resource to design around. The result connects an implementable coordination rule with a mathematical account of how much learning performance it preserves.

Scope: The agents learn in independent copies of the same finite-state environment. These results do not establish the same benefits for heterogeneous tasks, deep RL, or LLM-agent teams. Fewer synchronization rounds do not directly measure bytes, energy savings, or wall-clock speed.

Technical context & sources

The setup assumes bounded rewards, finite state and action spaces, and a finite-diameter Markov decision process. The coordinator computes an optimistic policy. The analysis assumes agents can receive a synchronization signal immediately and stop the current epoch.

The regret bound is Õ(DS√(MAT)), and synchronization is O(MSA log(MT)): M agents, S states, A actions, horizon T, and diameter D. The comparator, MOD-UCRL2, communicates each time step and processes agent requests sequentially. Matching this bound is not a claim of optimal dependence on every parameter.

Results plots are unmodified crops of Figures 1(a,c) and 2(a,c), page 8 of the supplied paper (proceedings page 254). Figures show averages over 50 independent runs; the regret plots also include the paper’s error bars. See Sections 3–7 and Theorems 1–3 for assumptions, algorithms, and analysis.