LAPSE:2026.1210
Published Article
LAPSE:2026.1210
Local-Global Learning of Interpretable Control Polices: The Interface between MPC and Reinforcement Learning
Ali Mesbah
July 13, 2026
Abstract
Optimal decision-making under uncertainty is a shared challenge across modern chemical, manufacturing, and energy systems that increasingly demand safe, data-driven autonomy. This talk revisits optimal control through the lens of the Bellman equation, emphasizing how optimal control theory and reinforcement learning have developed complementary, yet largely disconnected, perspectives on global optimality. In one view, central to reinforcement learning, the Bellman equation defines a global optimality condition that guides iterative policy learning from interacting with the system, but typically yields opaque control laws that are difficult to interpret, and deploy in safety-critical settings. In another view, widely adopted in model predictive control (MPC), the Bellman equation underpins tractable finite-horizon optimizations that deliver interpretable, constraint-aware, and modular local controllers, yet without explicit guarantees on alignment with global optimality. Building on these ideas, we introduce a local-global paradigm that treats MPC and related optimization-based controllers as structured function approximators designed to approximately satisfy the global Bellman optimality condition. We discuss algorithmic strategies for learning interpretable local decision makers whose adaptation is guided by Bellman residuals, along with the benefits and practical challenges that arise in terms of stability, constraint satisfaction, and sample efficiency. These concepts are illustrated through case studies that unify reinforcement learning and MPC for safe, high-performance control in complex, uncertain dynamical systems. The talk concludes by outlining open problems and research opportunities in learning interpretable control policies that achieve globally optimal performance while retaining the transparency and reliability required for real-world process control and optimization applications.
Suggested Citation
Mesbah A. Local-Global Learning of Interpretable Control Polices: The Interface between MPC and Reinforcement Learning. (2026). LAPSE:2026.1210
Author Affiliations
Mesbah A: University of California, Berkeley, Department of Chemical and Biomolecular Engineering
Journal Name
Proceedings of FOPAM 2026
Volume
0
First Page
12
Last Page
12
Year
2026
Publication Date
2026-07-13
Version Comments
Original Submission
Other Meta
PII: 0012-0012-10-PSE-0-2026, Publication Type: Abstract
Record Map
Published Article

LAPSE:2026.1210
This Record
External Link

https://doi.org/10.69997/pse.112403
Publisher Version
Download
Files
Jul 13, 2026
Main Article
License
CC BY-SA 4.0
Meta
Record Statistics
Record Views
95
Version History
[v1] (Original Submission)
Jul 13, 2026
 
Verified by curator on
Jul 13, 2026
This Version Number
v1
Citations
Most Recent
This Version
URL Here
https://psecommunity.org/LAPSE:2026.1210
 
Record Owner
PSE Press
Links to Related Works
Directly Related to This Work
Publisher Version
(0.13 seconds)

[0.14 s]