Overview
This study focuses on over-the-air (OTA) computation enabled online federated learning (FL) within low-Earth orbit (LEO) satellite networks. The research addresses the optimization of a dual-layer OTA aggregation architecture. This architecture involves ground devices uploading analog model updates to serving satellites via an initial uplink OTA aggregation, followed by satellites forwarding these aggregated signals to a data processing center through a second OTA aggregation round.
Research Context
The problem investigated is the maximization of long-term data utilization. This objective considers scenarios where ground devices continuously generate new data, and previously collected samples progressively diminish in freshness. The optimization is constrained by specific operational parameters, including the satellite beam budget, the transmit-power limit of devices, and a global mean squared error (MSE) constraint. This MSE constraint is designed to manage the end-to-end distortion during aggregation.
Approach
The core problem formulation resulted in a coupled mixed-integer nonlinear programming (MINLP) problem. This formulation encompasses tightly linked discrete beam-hopping decisions and continuous power control mechanisms. The combinatorial nature of the action space and the presence of nonconvex constraints rendered the problem NP-hard and computationally intractable. Furthermore, the dynamic nature of LEO satellite networks, characterized by time-varying satellite topology and continuous data generation, framed the problem as a sequential decision-making challenge requiring adaptive online scheduling.
To address these complexities, the research reframed the problem as a Markov decision process. A deep reinforcement learning framework, specifically based on proximal policy optimization (PPO), was then developed. This framework was designed to jointly optimize both adaptive beam hopping and power control. An MSE-aware reward function was incorporated within the framework to balance data utilization and aggregation accuracy.
Findings
- The proposed algorithm consistently outperformed other benchmark schemes.
- The algorithm achieved superior long-term data utilization.
- The algorithm demonstrated faster federated learning convergence.
- The algorithm satisfied the specified mean squared error (MSE) requirement.
Why This Matters
The findings indicate a method for enhancing the efficiency and performance of federated learning in dynamic LEO satellite environments. By jointly optimizing beam hopping and power control, the approach addresses challenges related to data freshness and aggregation distortion, which are pertinent for maintaining the efficacy of distributed machine learning in space-based networks.