LAPSE:2026.0521v1
Published Article

LAPSE:2026.0521v1
Relating Loss Geometry to Empirical Generalization in Recurrent Neural Net Surrogates: Three Tanks Case Study
June 12, 2026
Abstract
Recurrent neural nets (RNNs) are now commonly used for the surrogate modeling of process systems, leading to better control and faster real-time optimization. However, when trained with small training data sets, most experiments show that RNNs exhibit poor generalization abilities outside the range of the training data space. Nonetheless, recent advances in deep learning research have shown that certain characteristics of the loss landscape of trained models, such as the flatness around the local minimum, tend to relate to generalization ability. This paper investigates this phenomenon for the case of RNN surrogates of the well-known Three Tanks case study, which is representative of many continuous processes. We trained a total of 200 LSTMs (long short-term memory networks) differing in initialization, architecture, and training dynamics on the same data of 500 samples. The number of model parameters ranges from 238 to 11, 353. We estimated the loss curvature of each trained model using the Hessian-vector products method and then evaluated their test accuracy on 100 pre-generated data sets with 50% larger input amplitude than the training data to force extrapolation. We found that the estimate of the top eigenvalue of the Hessian matrix is 75% rank-correlated to the one-step-ahead prediction accuracy averaged on all test data sets. We report other insights on the impact of other Hessian-based metrics to generalization ability. Our study can potentially inform how to control the training dynamics of RNN surrogates to improve their extrapolation ability in the absence of first-principles knowledge.
Recurrent neural nets (RNNs) are now commonly used for the surrogate modeling of process systems, leading to better control and faster real-time optimization. However, when trained with small training data sets, most experiments show that RNNs exhibit poor generalization abilities outside the range of the training data space. Nonetheless, recent advances in deep learning research have shown that certain characteristics of the loss landscape of trained models, such as the flatness around the local minimum, tend to relate to generalization ability. This paper investigates this phenomenon for the case of RNN surrogates of the well-known Three Tanks case study, which is representative of many continuous processes. We trained a total of 200 LSTMs (long short-term memory networks) differing in initialization, architecture, and training dynamics on the same data of 500 samples. The number of model parameters ranges from 238 to 11, 353. We estimated the loss curvature of each trained model using the Hessian-vector products method and then evaluated their test accuracy on 100 pre-generated data sets with 50% larger input amplitude than the training data to force extrapolation. We found that the estimate of the top eigenvalue of the Hessian matrix is 75% rank-correlated to the one-step-ahead prediction accuracy averaged on all test data sets. We report other insights on the impact of other Hessian-based metrics to generalization ability. Our study can potentially inform how to control the training dynamics of RNN surrogates to improve their extrapolation ability in the absence of first-principles knowledge.
Record ID
Keywords
Artificial Intelligence, Derivative Free Optimization, Dynamic Modelling, Generalization, Hessian vector products, Machine Learning, System Identification
Subject
Suggested Citation
Roxas RM II, Pilario KE. Relating Loss Geometry to Empirical Generalization in Recurrent Neural Net Surrogates: Three Tanks Case Study. Systems and Control Transactions 5:2542-2550 (2026) https://doi.org/10.69997/sct.112163
Author Affiliations
Roxas RM II: Process Systems Engineering Laboratory, Department of Chemical Engineering, University of the Philippines Diliman, Quezon City, 1101, Philippines [ORCID]
Pilario KE: Process Systems Engineering Laboratory, Department of Chemical Engineering, University of the Philippines Diliman, Quezon City, 1101, Philippines [ORCID]
[Login] to see author email addresses.
Pilario KE: Process Systems Engineering Laboratory, Department of Chemical Engineering, University of the Philippines Diliman, Quezon City, 1101, Philippines [ORCID]
[Login] to see author email addresses.
Journal Name
Systems and Control Transactions
Volume
5
First Page
2542
Last Page
2550
Year
2026
Publication Date
2026-06-12
Version Comments
Original Submission
Other Meta
PII: 2542-2550-469-SCT-5-2026, Publication Type: Journal Article
Record Map
Published Article

LAPSE:2026.0521v1
This Record
External Link

https://doi.org/10.69997/sct.112163
Publisher Version
Download
Meta
Record Statistics
Record Views
298
Version History
[v1] (Original Submission)
Jun 12, 2026
Verified by curator on
Jun 12, 2026
This Version Number
v1
Citations
Most Recent
This Version
URL Here
http://psecommunity.org/LAPSE:2026.0521v1
Record Owner
PSE Press
Links to Related Works
References Cited
- Pilario KES, Cao Y, Shafiee M. A kernel design approach to improve kernel subspace identification. IEEE Trans. Ind. Electron. 68:6171-6180 (2021) https://doi.org/10.1109/tie.2020.2996142
- Wu Z, Christofides PD, Wu W, Wang Y, Abdullah F, Alnajdi A, Kadakia Y. A tutorial review of machine learning-based model predictive control methods. Reviews in Chemical Engineering 41:359-400 (2024) https://doi.org/10.1515/revce-2024-0055
- Sun W, Paiva ARC, Xu P, Sundaram A, Braatz RD. Fault detection and identification using bayesian recurrent neural networks. Computers & Chemical Engineering 141:106991 (2020) https://doi.org/10.1016/j.compchemeng.2020.106991
- Chen H, Cen J, Yang Z, Si W, Cheng H. Fault diagnosis of the dynamic chemical process based on the optimized CNN-LSTM network. ACS Omega 7:34389-34400 (2022) https://doi.org/10.1021/acsomega.2c04017
- Zhu W, Chebeir J, Webb Z, Romagnoli J. A Deep Learning Approach on Surrogate Model Optimization of a Cryogenic NGL Recovery Unit Operation. Computer Aided Chemical Engineering, 48:1285-1290 (2020)
- Esche E, Weigert J, Brand Rihm G, Göbel J, Repke JU. Architectures for neural networks as surrogates for dynamic systems in chemical engineering. Chemical Engineering Research and Design 177:184-199 (2022) https://doi.org/10.1016/j.cherd.2021.10.042
- Alhajeri MS, Alnajdi A, Abdullah F, Christofides PD. On generalization error of neural network models and its application to predictive control of nonlinear processes. Chem. Eng. Res. Des. 189:664-679 (2023) https://doi.org.10.1016/j.cherd.2022.12.001
- Cheng X, Huang K, Ma S. Generalization and risk bounds for recurrent neural networks. Neurocomputing 616:128825 (2025) https://doi.org/10.1016/j.neucom.2024.128825
- Li J, Qin SJ. Applying and dissecting LSTM neural networks and regularized learning for dynamic inferential modeling. Computers & Chemical Engineering 175:108264 (2023) https://doi.org/10.1016/j.compchemeng.2023.108264
- ?awry?czuk M. Input convex neural networks in nonlinear predictive control: a multi-model approach. Neurocomputing 513:273-293 (2022) https://doi.org/10.1016/j.neucom.2022.09.108
- Pravin PS, Tan JZM, Yap KS, Wu Z. Hyperparameter optimization strategies for machine learning-based stochastic energy efficient scheduling in cyber-physical production systems. Digital Chemical Engineering 4:100047 (2022) https://doi.org/10.1016/j.dche.2022.100047
- Shakouri B, Murthy S, De Clippeleir H, Lesnik K, Massoudieh A. Surrogate modeling of ASM1 using feedforward and LSTM neural networks. Journal of Water Process Engineering 81:109246 (2026) https://doi.org/10.1016/j.jwpe.2025.109246
- Hochreiter S, Schmidhuber J. Flat minima. Neural Computation 9:1-42 (1997) https://doi.org/10.1162/neco.1997.9.1.1
- Chaudhari P, Choromanska A, Soatto S, LeCun Y, Baldassi C, Borgs C, Chayes J, Sagun L, Zecchina R. Entropy-SGD: Biasing Gradient Descent Into Wide Valleys. (2017) http://arxiv.org/abs/1611.01838
- Dinh L, Pascanu R, Bengio S, Bengio Y. Sharp Minima Can Generalize For Deep Nets. (2017) http://arxiv.org/abs/1703.04933
- Keskar NS, Mudigere D, Nocedal J, Smelyanskiy M, Tang PTP. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. (2017) http://arxiv.org/abs/1609.04836
- Tsuzuku Y, Sato I, Sugiyama M. Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks Using PAC-Bayesian Analysis. Proceedings of the 37th International Conference on Machine Learning 119 (2020)
- Wilson AG. Deep Learning is Not So Mysterious or Different. (2025) http://arxiv.org/abs/2503.02113
- Shoham N, Mor-Yosef L, Avron H. Flatness After All? (2025) http://arxiv.org/abs/2506.17809
- Foret P, Kleiner A, Mobahi H, Neyshabur B. Sharpness-Aware Minimization for Efficiently Improving Generalization. (2021) http://arxiv.org/abs/2010.01412
- Zhang X, Xu R, Yu H, Dong Y, Tian P, Cu P. Flatness-Aware Minimization for Domain Generalization. (2023) http://arxiv.org/abs/2307.11108
- Lee S, He C, Avestimehr S. Achieving small-batch accuracy with large-batch scalability via hessian-aware learning rate adjustment. Neural Networks 158:1-14 (2023) https://doi.org/10.1016/j.neunet.2022.11.007
- Yoshida Y, Miyato T. Spectral Norm Regularization for Improving the Generalizability of Deep Learning. (2017) http://arxiv.org/abs/1705.10941
- Saab S Jr, Fu Y, Ray A, Hauser M. A dynamically stabilized recurrent neural network. Neural Process Lett 54:1195-1209 (2021) https://doi.org/10.1007/s11063-021-10676-7
- Pilario KE, Wu Z. Fast mixed kernel canonical variate analysis for learning-based nonlinear model predictive control. Chemical Engineering Research and Design 219:19-33 (2025) https://doi.org/10.1016/j.cherd.2025.05.050
- Pearlmutter BA. Fast exact multiplication by the hessian. Neural Computation 6:147-160 (1994) https://doi.org/10.1162/neco.1994.6.1.147
- Yao Z, Gholami A, Keutzer K, Mahoney MW. Pyhessian: neural networks through the lens of the hessian. 2020 IEEE International Conference on Big Data (Big Data) :581-590 (2020) https://doi.org/10.1109/bigdata50022.2020.9378171
(0.14 seconds)
[0.15 s]

