Proceedings of ESCAPE 36ISSN: 2818-4734
Volume: 5 (2026)
Table of Contents
LAPSE:2026.0521
Published Article
LAPSE:2026.0521
Relating Loss Geometry to Empirical Generalization in Recurrent Neural Net Surrogates: Three Tanks Case Study
June 12, 2026
Abstract
Recurrent neural nets (RNNs) are now commonly used for the surrogate modeling of process systems, leading to better control and faster real-time optimization. However, when trained with small training data sets, most experiments show that RNNs exhibit poor generalization abilities outside the range of the training data space. Nonetheless, recent advances in deep learning research have shown that certain characteristics of the loss landscape of trained models, such as the flatness around the local minimum, tend to relate to generalization ability. This paper investigates this phenomenon for the case of RNN surrogates of the well-known Three Tanks case study, which is representative of many continuous processes. We trained a total of 200 LSTMs (long short-term memory networks) differing in initialization, architecture, and training dynamics on the same data of 500 samples. The number of model parameters ranges from 238 to 11, 353. We estimated the loss curvature of each trained model using the Hessian-vector products method and then evaluated their test accuracy on 100 pre-generated data sets with 50% larger input amplitude than the training data to force extrapolation. We found that the estimate of the top eigenvalue of the Hessian matrix is 75% rank-correlated to the one-step-ahead prediction accuracy averaged on all test data sets. We report other insights on the impact of other Hessian-based metrics to generalization ability. Our study can potentially inform how to control the training dynamics of RNN surrogates to improve their extrapolation ability in the absence of first-principles knowledge.
Keywords
Artificial Intelligence, Derivative Free Optimization, Dynamic Modelling, Generalization, Hessian vector products, Machine Learning, System Identification
Suggested Citation
Roxas RM II, Pilario KE. Relating Loss Geometry to Empirical Generalization in Recurrent Neural Net Surrogates: Three Tanks Case Study. Systems and Control Transactions 5:2542-2550 (2026) https://doi.org/10.69997/sct.112163
Author Affiliations
Roxas RM II: Process Systems Engineering Laboratory, Department of Chemical Engineering, University of the Philippines Diliman, Quezon City, 1101, Philippines [ORCID]
Pilario KE: Process Systems Engineering Laboratory, Department of Chemical Engineering, University of the Philippines Diliman, Quezon City, 1101, Philippines [ORCID]
[Login] to see author email addresses.
Journal Name
Systems and Control Transactions
Volume
5
First Page
2542
Last Page
2550
Year
2026
Publication Date
2026-06-12
Version Comments
Original Submission
Other Meta
PII: 2542-2550-469-SCT-5-2026, Publication Type: Journal Article
Record Map
Published Article

LAPSE:2026.0521
This Record
External Link

https://doi.org/10.69997/sct.112163
Publisher Version
Download
Files
Jun 12, 2026
Main Article
License
CC BY-SA 4.0
Meta
Record Statistics
Record Views
252
Version History
[v1] (Original Submission)
Jun 12, 2026
 
Verified by curator on
Jun 12, 2026
This Version Number
v1
Citations
Most Recent
This Version
URL Here
http://psecommunity.org/LAPSE:2026.0521
 
Record Owner
PSE Press
Links to Related Works
Directly Related to This Work
Publisher Version
References Cited
  1. Pilario KES, Cao Y, Shafiee M. A kernel design approach to improve kernel subspace identification. IEEE Trans. Ind. Electron. 68:6171-6180 (2021) https://doi.org/10.1109/tie.2020.2996142
  2. Wu Z, Christofides PD, Wu W, Wang Y, Abdullah F, Alnajdi A, Kadakia Y. A tutorial review of machine learning-based model predictive control methods. Reviews in Chemical Engineering 41:359-400 (2024) https://doi.org/10.1515/revce-2024-0055
  3. Sun W, Paiva ARC, Xu P, Sundaram A, Braatz RD. Fault detection and identification using bayesian recurrent neural networks. Computers & Chemical Engineering 141:106991 (2020) https://doi.org/10.1016/j.compchemeng.2020.106991
  4. Chen H, Cen J, Yang Z, Si W, Cheng H. Fault diagnosis of the dynamic chemical process based on the optimized CNN-LSTM network. ACS Omega 7:34389-34400 (2022) https://doi.org/10.1021/acsomega.2c04017
  5. Zhu W, Chebeir J, Webb Z, Romagnoli J. A Deep Learning Approach on Surrogate Model Optimization of a Cryogenic NGL Recovery Unit Operation. Computer Aided Chemical Engineering, 48:1285-1290 (2020)
  6. Esche E, Weigert J, Brand Rihm G, Göbel J, Repke JU. Architectures for neural networks as surrogates for dynamic systems in chemical engineering. Chemical Engineering Research and Design 177:184-199 (2022) https://doi.org/10.1016/j.cherd.2021.10.042
  7. Alhajeri MS, Alnajdi A, Abdullah F, Christofides PD. On generalization error of neural network models and its application to predictive control of nonlinear processes. Chem. Eng. Res. Des. 189:664-679 (2023) https://doi.org.10.1016/j.cherd.2022.12.001
  8. Cheng X, Huang K, Ma S. Generalization and risk bounds for recurrent neural networks. Neurocomputing 616:128825 (2025) https://doi.org/10.1016/j.neucom.2024.128825
  9. Li J, Qin SJ. Applying and dissecting LSTM neural networks and regularized learning for dynamic inferential modeling. Computers & Chemical Engineering 175:108264 (2023) https://doi.org/10.1016/j.compchemeng.2023.108264
  10. ?awry?czuk M. Input convex neural networks in nonlinear predictive control: a multi-model approach. Neurocomputing 513:273-293 (2022) https://doi.org/10.1016/j.neucom.2022.09.108
  11. Pravin PS, Tan JZM, Yap KS, Wu Z. Hyperparameter optimization strategies for machine learning-based stochastic energy efficient scheduling in cyber-physical production systems. Digital Chemical Engineering 4:100047 (2022) https://doi.org/10.1016/j.dche.2022.100047
  12. Shakouri B, Murthy S, De Clippeleir H, Lesnik K, Massoudieh A. Surrogate modeling of ASM1 using feedforward and LSTM neural networks. Journal of Water Process Engineering 81:109246 (2026) https://doi.org/10.1016/j.jwpe.2025.109246
  13. Hochreiter S, Schmidhuber J. Flat minima. Neural Computation 9:1-42 (1997) https://doi.org/10.1162/neco.1997.9.1.1
  14. Chaudhari P, Choromanska A, Soatto S, LeCun Y, Baldassi C, Borgs C, Chayes J, Sagun L, Zecchina R. Entropy-SGD: Biasing Gradient Descent Into Wide Valleys. (2017) http://arxiv.org/abs/1611.01838
  15. Dinh L, Pascanu R, Bengio S, Bengio Y. Sharp Minima Can Generalize For Deep Nets. (2017) http://arxiv.org/abs/1703.04933
  16. Keskar NS, Mudigere D, Nocedal J, Smelyanskiy M, Tang PTP. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. (2017) http://arxiv.org/abs/1609.04836
  17. Tsuzuku Y, Sato I, Sugiyama M. Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks Using PAC-Bayesian Analysis. Proceedings of the 37th International Conference on Machine Learning 119 (2020)
  18. Wilson AG. Deep Learning is Not So Mysterious or Different. (2025) http://arxiv.org/abs/2503.02113
  19. Shoham N, Mor-Yosef L, Avron H. Flatness After All? (2025) http://arxiv.org/abs/2506.17809
  20. Foret P, Kleiner A, Mobahi H, Neyshabur B. Sharpness-Aware Minimization for Efficiently Improving Generalization. (2021) http://arxiv.org/abs/2010.01412
  21. Zhang X, Xu R, Yu H, Dong Y, Tian P, Cu P. Flatness-Aware Minimization for Domain Generalization. (2023) http://arxiv.org/abs/2307.11108
  22. Lee S, He C, Avestimehr S. Achieving small-batch accuracy with large-batch scalability via hessian-aware learning rate adjustment. Neural Networks 158:1-14 (2023) https://doi.org/10.1016/j.neunet.2022.11.007
  23. Yoshida Y, Miyato T. Spectral Norm Regularization for Improving the Generalizability of Deep Learning. (2017) http://arxiv.org/abs/1705.10941
  24. Saab S Jr, Fu Y, Ray A, Hauser M. A dynamically stabilized recurrent neural network. Neural Process Lett 54:1195-1209 (2021) https://doi.org/10.1007/s11063-021-10676-7
  25. Pilario KE, Wu Z. Fast mixed kernel canonical variate analysis for learning-based nonlinear model predictive control. Chemical Engineering Research and Design 219:19-33 (2025) https://doi.org/10.1016/j.cherd.2025.05.050
  26. Pearlmutter BA. Fast exact multiplication by the hessian. Neural Computation 6:147-160 (1994) https://doi.org/10.1162/neco.1994.6.1.147
  27. Yao Z, Gholami A, Keutzer K, Mahoney MW. Pyhessian: neural networks through the lens of the hessian. 2020 IEEE International Conference on Big Data (Big Data) :581-590 (2020) https://doi.org/10.1109/bigdata50022.2020.9378171
(0.1 seconds)

[0.1 s]