LAPSE:2026.0399
Published Article

LAPSE:2026.0399
Control-Guided Reinforcement Learning for Cooperative Energy Management
June 12, 2026
Abstract
Addressing the urgent transition to low-carbon energy systems requires microgrids capable of locally coordinating electricity generation, storage, and flexible consumption. Their efficient integration calls for control strategies that are scalable, privacy-preserving, and robust to uncertainty. To address such a challenging control problem, this work proposes a decentralised Multi-Agent Reinforcement Learning (MARL) approach based on the Cross-Entropy Method (CEM) for the coordination of prosumers, equipped with renewable generation and vehicle-to-grid capabilities. To improve sample efficiency and robustness, the policy is warm-started using Behaviour Cloning (BC) from a classical Proportional-Integral-Derivative (PID) controller, resulting in a hybrid BC-CEM framework. The proposed method is evaluated in a realistic microgrid simulation with stochastic demand and real weather and generation profiles. Results show that BC-CEM accelerates convergence and achieves lower energy costs compared to both PID control and randomly initialized CEM, without sacrificing comfort or mobility requirements. The findings highlight the effectiveness of combining derivative-free optimization with imitation learning in complex MARL tasks, such as energy flexibility coordination.
Addressing the urgent transition to low-carbon energy systems requires microgrids capable of locally coordinating electricity generation, storage, and flexible consumption. Their efficient integration calls for control strategies that are scalable, privacy-preserving, and robust to uncertainty. To address such a challenging control problem, this work proposes a decentralised Multi-Agent Reinforcement Learning (MARL) approach based on the Cross-Entropy Method (CEM) for the coordination of prosumers, equipped with renewable generation and vehicle-to-grid capabilities. To improve sample efficiency and robustness, the policy is warm-started using Behaviour Cloning (BC) from a classical Proportional-Integral-Derivative (PID) controller, resulting in a hybrid BC-CEM framework. The proposed method is evaluated in a realistic microgrid simulation with stochastic demand and real weather and generation profiles. Results show that BC-CEM accelerates convergence and achieves lower energy costs compared to both PID control and randomly initialized CEM, without sacrificing comfort or mobility requirements. The findings highlight the effectiveness of combining derivative-free optimization with imitation learning in complex MARL tasks, such as energy flexibility coordination.
Record ID
Keywords
Behavioral Cloning, Derivative-Free Optimization, Energy Management, Machine Learning, Reinforcement Learning
Subject
Suggested Citation
Moreno-Palancas IF, Díaz RS, Femenía RR, Caballero JA, Chanona ADR. Control-Guided Reinforcement Learning for Cooperative Energy Management. Systems and Control Transactions 5:1558-1564 (2026) https://doi.org/10.69997/sct.165580
Author Affiliations
Moreno-Palancas IF: University of Alicante, Department of Chemical Engineering, Alicante, Comunidad Valenciana, Spain [ORCID]
Díaz RS: University of Alicante, Department of Chemical Engineering, Alicante, Comunidad Valenciana, Spain [ORCID]
Femenía RR: University of Alicante, Department of Chemical Engineering, Alicante, Comunidad Valenciana, Spain [ORCID]
Caballero JA: University of Alicante, Department of Chemical Engineering, Alicante, Comunidad Valenciana, Spain [ORCID]
Chanona ADR: Imperial College London, Sargent Centre for Process Systems Engineering, Department of Chemical Engineering, London, United Kingdom [ORCID]
[Login] to see author email addresses.
Díaz RS: University of Alicante, Department of Chemical Engineering, Alicante, Comunidad Valenciana, Spain [ORCID]
Femenía RR: University of Alicante, Department of Chemical Engineering, Alicante, Comunidad Valenciana, Spain [ORCID]
Caballero JA: University of Alicante, Department of Chemical Engineering, Alicante, Comunidad Valenciana, Spain [ORCID]
Chanona ADR: Imperial College London, Sargent Centre for Process Systems Engineering, Department of Chemical Engineering, London, United Kingdom [ORCID]
[Login] to see author email addresses.
Journal Name
Systems and Control Transactions
Volume
5
First Page
1558
Last Page
1564
Year
2026
Publication Date
2026-06-12
Version Comments
Original Submission
Other Meta
PII: 1558-1564-100-SCT-5-2026, Publication Type: Journal Article
Record Map
Published Article

LAPSE:2026.0399
This Record
External Link

https://doi.org/10.69997/sct.165580
Publisher Version
Conference Presentation

LAPSE:2026.0610
Control-Guided Reinforcement Learni...
Download
Meta
Record Statistics
Record Views
276
Version History
[v1] (Original Submission)
Jun 12, 2026
Verified by curator on
Jun 12, 2026
This Version Number
v1
Citations
Most Recent
This Version
URL Here
https://psecommunity.org/LAPSE:2026.0399
Record Owner
PSE Press
Links to Related Works
LAPSE Records Linking Here
Directly Related to This Work
Control-Guided Reinforcement Learning for Cooperative Energy Management
References Cited
- IEA. Unlocking the Potential of Distributed Energy Resources. IEA (2022) Licence: CC by 4.0 https://www.iea.org/reports/unlocking-the-potential-of-distributed-energy-resources
- Mohammadi P, Darshi R, Shamaghdari S, Siano P. Comparative analysis of control strategies for microgrid energy management with a focus on reinforcement learning. IEEE Access 12:171368-171395 (2024) https://doi.org/10.1109/access.2024.3495032
- Zhang H, Seal S, Wu , Bouffard F, Boulet B. Building energy management with reinforcement learning and model predictive control: a survey. IEEE Access 10:27853-27862 (2022) https://doi.org/10.1109/access.2022.3156581
- Zhang B, Hu W, MYM Ghias A, Xu X, Chen Z. Multi-agent deep reinforcement learning based distributed control architecture for interconnected multi-energy microgrid energy management and optimization. Energy Conversion and Management 277:116647 (2023) https://doi.org/10.1016/j.enconman.2022.116647
- Charbonnier F, Morstyn T, McCulloch MD. Scalable multi-agent reinforcement learning for distributed control of residential energy flexibility. Applied Energy 314:118825 (2022) https://doi.org/10.1016/j.apenergy.2022.118825
- Han Y, Wu J, Chen H, Si F, Cao Z, Zhao Q. Enhancing grid-interactive buildings demand response: sequential update-based multiagent deep reinforcement learning approach. IEEE Internet Things J. 11:24439-24451 (2024) https://doi.org/10.1109/jiot.2024.3357109
- Charbonnier F, Peng B, Vienne J, Stai E, Morstyn T, McCulloch M. Centralised rehearsal of decentralised cooperation: multi-agent reinforcement learning for the scalable coordination of residential energy flexibility. Applied Energy 377:124406 (2025) https://doi.org/10.1016/j.apenergy.2024.124406
- Savino S, Minella T, Nagy Z, Capozzoli A. A scalable demand-side energy management control strategy for large residential districts based on an attention-driven multi-agent DRL approach. Applied Energy 393:125993 (2025) https://doi.org/10.1016/j.apenergy.2025.125993
- Ji C, Xiao H, Pei W, Wang X, Pu X. Coordinated voltage control for distribution network and multi-microgrids based on improved EM-RACE multi-agent reinforcement learning. Int. J. Electr. Power Energy Syst. 172:111315 (2025) https://doi.org/10.1016/j.ijepes.2025.111315
- Bloor M, Ahmed A, Kotecha N, Mercangöz M, Tsay C, del Río-Chanona EA. Control-informed reinforcement learning for chemical processes. Ind. Eng. Chem. Res. 64:4966-4978 (2025) https://doi.org/10.1021/acs.iecr.4c03233
- Salimans T, Ho J, Chen X, Sidor S, Sutskever I. Evolution Strategies as a Scalable Alternative to Reinforcement Learning. arXiv (2017). https://arxiv.org/abs/1703.03864
- de Boer PT, Kroese DP, Mannor S, Rubinstein RY. A tutorial on the cross-entropy method. Ann Oper Res 134:19-67 (2005) https://doi.org/10.1007/s10479-005-5724-z
- Zhang Z, Jin J, Jagersand M, Luo J, Schuurmans D. A Simple Decentralized Cross-Entropy Method. arXiv (2022) https://doi.org/10.48550/arXiv.2212.08235
- Levine S, Kumar A, Tucker G, Fu J. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv (2020) https://doi.org/10.48550/arXiv.2005.01643
(0.14 seconds)
[0.15 s]

