← All research notes

Edge learning / Distributed optimization

Where should distributed learning happen?

CE-FL: share the training work—and move the server that brings it together.

The takeaway

Distributed learning has a placement problem: where should training happen, and where should updates meet? CE-FL adapts both to the data, devices, and network.

01 / The network is part of the problem

A fleet of connected devices can have very different amounts of data, computing power, and connectivity. Asking every device to train locally and always send updates to the same server can waste time and energy.

Devices connect through base stations to multiple edge servers
Learning across three layers. Devices and edge servers share the training work; base stations connect them. CE-FL can offload some training data as well as exchange model updates. Figure 2, page 2.

02 / Let the meeting point move

CE-FL jointly decides how much training happens at each location, how data travels, and which edge server combines the model updates. This “floating” aggregation point can change from round to round.

Aggregation server choices change across rounds for CE-FL and two greedy baselines
No single server wins every round. The red CE-FL path differs from rules that only favor the most data or the best data rate. Server names label the modeled network locations. Figure 3(c), page 14.

03 / Does moving it help?

The paper evaluates a modeled network of 20 devices, 10 base stations, and 5 servers, informed by measurements from 5G/4G and CBRS testbeds. These plots isolate the choice of aggregation server.

Lower bars are better

Compare CE-FL with a fixed server and two greedy selection rules. The CE-FL color differs between plots: orange for delay, blue for energy.

Global aggregation delay for CE-FL and server-selection baselines
Less waiting. CE-FL has lower aggregation delay in the five rounds shown. Figure 4(a).
Average processing-unit energy for CE-FL and server-selection baselines
Less energy. Jointly considering network conditions reduces average processing-unit energy here. Figure 4(b).

Select any figure for its full-size view.

39.9% less energy

In a separate end-to-end comparison at 80% Fashion-MNIST accuracy, CE-FL used 47.3 kJ versus FedNova’s 78.7 kJ, with 15.7% lower training delay. These are results for the evaluated setup, not universal savings. Tables I–II.

Why this work matters

It treats model training and network resource allocation as one connected problem. The contribution combines a cooperative learning architecture, convergence analysis, and a distributed optimization method for deciding where the work goes.

Scope: These are numerical evaluations informed by testbed measurements. Data offloading requires permission to move the relevant data; CE-FL does not guarantee that all raw data stays on devices.

Technical context & sources

The resource decisions include data routing, local SGD iterations, mini-batch sizes, processing rates, and the aggregation server. The paper accounts for changing local datasets and analyzes non-convex learning and a distributed optimization solver. Its solver guarantee concerns a stationary solution under the stated assumptions, not a globally optimal solution.

Figures are unmodified crops of Figures 2, 3(c), and 4(a–b) from the supplied paper. Section VI and Appendices F–G describe the measurements and simulation setup. Tables I–II compare energy and delay at matched target accuracies on Fashion-MNIST and CIFAR-10; the highlighted result uses the 80% Fashion-MNIST column.