Edge learning / Distributed optimization
Where should distributed learning happen?
CE-FL: share the training work—and move the server that brings it together.
Distributed learning has a placement problem: where should training happen, and where should updates meet? CE-FL adapts both to the data, devices, and network.
01 / The network is part of the problem
A fleet of connected devices can have very different amounts of data, computing power, and connectivity. Asking every device to train locally and always send updates to the same server can waste time and energy.

02 / Let the meeting point move
CE-FL jointly decides how much training happens at each location, how data travels, and which edge server combines the model updates. This “floating” aggregation point can change from round to round.

03 / Does moving it help?
The paper evaluates a modeled network of 20 devices, 10 base stations, and 5 servers, informed by measurements from 5G/4G and CBRS testbeds. These plots isolate the choice of aggregation server.
Compare CE-FL with a fixed server and two greedy selection rules. The CE-FL color differs between plots: orange for delay, blue for energy.


Select any figure for its full-size view.
In a separate end-to-end comparison at 80% Fashion-MNIST accuracy, CE-FL used 47.3 kJ versus FedNova’s 78.7 kJ, with 15.7% lower training delay. These are results for the evaluated setup, not universal savings. Tables I–II.
Why this work matters
It treats model training and network resource allocation as one connected problem. The contribution combines a cooperative learning architecture, convergence analysis, and a distributed optimization method for deciding where the work goes.
Scope: These are numerical evaluations informed by testbed measurements. Data offloading requires permission to move the relevant data; CE-FL does not guarantee that all raw data stays on devices.
Technical context & sources
The resource decisions include data routing, local SGD iterations, mini-batch sizes, processing rates, and the aggregation server. The paper accounts for changing local datasets and analyzes non-convex learning and a distributed optimization solver. Its solver guarantee concerns a stationary solution under the stated assumptions, not a globally optimal solution.
Figures are unmodified crops of Figures 2, 3(c), and 4(a–b) from the supplied paper. Section VI and Appendices F–G describe the measurements and simulation setup. Tables I–II compare energy and delay at matched target accuracies on Fashion-MNIST and CIFAR-10; the highlighted result uses the 80% Fashion-MNIST column.