LAPSE:2026.0448v1
Published Article

LAPSE:2026.0448v1
Deep Kernel Learning with Kolmogorov-Arnold Networks for Bayesian Optimization
June 12, 2026
Abstract
Deep Kernel Learning (DKL) has emerged as a powerful framework for Bayesian Optimization (BO), via combining expressive representation learning models with typical Gaussian Processes (GPs) surrogate models. However, conventional DKL typically relies on weight-based feature extractors (e.g., multilayer perceptrons (MLPs)), which often lack interpretability and may suffer from overfitting under data scarcity or training instability, potentially leading to a degraded uncertainty quantification in GP models. Grounded in the Kolmogorov-Arnold representation theorem, this paper proposes a novel DKL-KAN framework that employs Kolmogorov-Arnold Networks (KANs) as adaptive feature extractors, formulating a DKL-KAN surrogate model. Unlike MLPs, the KANs learn data-driven univariate functions, yielding more sample-efficient and stable representations for regression under limited data regimes. Followed by the GP, the DKL-KAN facilitates end-to-end learning of expressive latent representations while maintaining robust uncertainty quantification. The proposed DKL-KAN framework is further validated with the Williams-Otto (W-O) process benchmarks. Compared to classical GPs and DKL-MLPs, the DKL-KANs consistently exhibit superior predictive accuracy and uncertainty calibration even under high-dimensional and scarce dataset. The results show that the KANs generate more structured and informative latent embeddings, thus enhancing the ability to capture complex physical nonlinearities in sparsely sampled edge regions of the design space where traditional models often falter. Moreover, the DKL-KAN is embedded into the BO and yields the efficient optimization performance on the W-O process benchmark, achieving faster early-stage improvement and converging to higher objective values. These results demonstrate the potential of DKL-KAN surrogates for accelerating BO in chemical engineering processes.
Deep Kernel Learning (DKL) has emerged as a powerful framework for Bayesian Optimization (BO), via combining expressive representation learning models with typical Gaussian Processes (GPs) surrogate models. However, conventional DKL typically relies on weight-based feature extractors (e.g., multilayer perceptrons (MLPs)), which often lack interpretability and may suffer from overfitting under data scarcity or training instability, potentially leading to a degraded uncertainty quantification in GP models. Grounded in the Kolmogorov-Arnold representation theorem, this paper proposes a novel DKL-KAN framework that employs Kolmogorov-Arnold Networks (KANs) as adaptive feature extractors, formulating a DKL-KAN surrogate model. Unlike MLPs, the KANs learn data-driven univariate functions, yielding more sample-efficient and stable representations for regression under limited data regimes. Followed by the GP, the DKL-KAN facilitates end-to-end learning of expressive latent representations while maintaining robust uncertainty quantification. The proposed DKL-KAN framework is further validated with the Williams-Otto (W-O) process benchmarks. Compared to classical GPs and DKL-MLPs, the DKL-KANs consistently exhibit superior predictive accuracy and uncertainty calibration even under high-dimensional and scarce dataset. The results show that the KANs generate more structured and informative latent embeddings, thus enhancing the ability to capture complex physical nonlinearities in sparsely sampled edge regions of the design space where traditional models often falter. Moreover, the DKL-KAN is embedded into the BO and yields the efficient optimization performance on the W-O process benchmark, achieving faster early-stage improvement and converging to higher objective values. These results demonstrate the potential of DKL-KAN surrogates for accelerating BO in chemical engineering processes.
Record ID
Keywords
Bayesian Optimization, Deep Kernel Learning, Kolmogorov-Arnold Network, Process Optimization
Subject
Suggested Citation
Shang Z, Yuan Z, Zhang L, Dai Y. Deep Kernel Learning with Kolmogorov-Arnold Networks for Bayesian Optimization. Systems and Control Transactions 5:1967-1973 (2026) https://doi.org/10.69997/sct.169600
Author Affiliations
Shang Z: School of Chemical Engineering, Sichuan University, Chengdu, 610065, P. R. China.
Yuan Z: Department of Chemical Engineering, Tsinghua University, Beijing, 100084, China.
Zhang L: Department of Chemical Engineering, The Sargent Centre for Process Systems Engineering, Imperial College London, London, SW7 2AZ, United Kingdom
Dai Y: School of Chemical Engineering, Sichuan University, Chengdu, 610065, P. R. China.
[Login] to see author email addresses.
Yuan Z: Department of Chemical Engineering, Tsinghua University, Beijing, 100084, China.
Zhang L: Department of Chemical Engineering, The Sargent Centre for Process Systems Engineering, Imperial College London, London, SW7 2AZ, United Kingdom
Dai Y: School of Chemical Engineering, Sichuan University, Chengdu, 610065, P. R. China.
[Login] to see author email addresses.
Journal Name
Systems and Control Transactions
Volume
5
First Page
1967
Last Page
1973
Year
2026
Publication Date
2026-06-12
Version Comments
Original Submission
Other Meta
PII: 1967-1973-107-SCT-5-2026, Publication Type: Journal Article
Record Map
Published Article

LAPSE:2026.0448v1
This Record
External Link

https://doi.org/10.69997/sct.169600
Publisher Version
Download
Meta
Record Statistics
Record Views
300
Version History
[v1] (Original Submission)
Jun 12, 2026
Verified by curator on
Jun 12, 2026
This Version Number
v1
Citations
Most Recent
This Version
URL Here
https://psecommunity.org/LAPSE:2026.0448v1
Record Owner
PSE Press
Links to Related Works
References Cited
- Zhou J, Shi T, Ren J, He C. Accelerating operation optimization of complex chemical processes: a novel framework integrating artificial neural network and mixed-integer linear programming. Chemical Engineering Journal 481:148421 (2024) https://doi.org/10.1016/j.cej.2023.148421
- Shields BJ, Stevens J, Li J, Parasram M, Damani F, Alvarado JIM, Janey JM, Adams RP, Doyle AG. Bayesian reaction optimization as a tool for chemical synthesis. Nature 590:89-96 (2021) https://doi.org/10.1038/s41586-021-03213-y
- Li H, Hao J, Qiao S. Ai?driven electrolyte additive selection to boost aqueous zn?ion batteries stability. Advanced Materials 36: (2024) https://doi.org/10.1002/adma.202411991
- S. Falkner, A. Klein, F. Hutter, BOHB: Robust and Efficient Hyperparameter Optimization at Scale, (2018) https://doi.org/10.48550/arXiv.1807.01774
- Winz J, Fromme F, Engell S. Bayesian optimization of gray-box process models using a modified upper confidence bound acquisition function. Computers & Chemical Engineering 194:108976 (2025) https://doi.org/10.1016/j.compchemeng.2024.108976
- A.G. Wilson, Z. Hu, R. Salakhutdinov, E.P. Xing, Deep Kernel Learning, in: Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, PMLR, 2016: pp. 370-378. https://proceedings.mlr.press/v51/wilson16.html (accessed October 10, 2025).
- I. Achituve, G. Chechik, E. Fetaya, Guided Deep Kernel Learning, (2023) https://doi.org/10.48550/arXiv.2302.09574
- Botteghi N, Guo M, Brune C. Deep kernel learning of dynamical models from high-dimensional noisy data. Sci Rep 12: (2022) https://doi.org/10.1038/s41598-022-25362-4
- Singh S, Hernández-Lobato JM. Deep kernel learning for reaction outcome prediction and optimization. Commun Chem 7: (2024) https://doi.org/10.1038/s42004-024-01219-x
- Gayon-Lombardo A, del Rio-Chanona EA, Pino-Muñoz CA, Brandon NP. Deep kernel bayesian optimisation for closed-loop electrode microstructure design with user-defined properties. Energy and AI 22:100608 (2025) https://doi.org/10.1016/j.egyai.2025.100608
- Z. Liu, Y. Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Soljacic, T.Y. Hou, M. Tegmark, KAN: Kolmogorov-Arnold Networks, (2025) https://doi.org/10.48550/arXiv.2404.19756
- S.A. Faroughi, F. Mostajeran, A.H. Mashhadzadeh, S. Faroughi, Scientific Machine Learning with Kolmogorov-Arnold Networks, (2025) https://doi.org/10.48550/arXiv.2507.22959
- Wang Y, Sun J, Bai J, Anitescu C, Eshaghi MS, Zhuang X, Rabczuk T, Liu Y. Kolmogorov-arnold-informed neural network: a physics-informed deep learning framework for solving forward and inverse problems based on kolmogorov-arnold networks. Computer Methods in Applied Mechanics and Engineering 433:117518 (2025) https://doi.org/10.1016/j.cma.2024.117518
- Zeng Y, Cao W, Zuo Y, Peng T, Hou Y, Miao L, Wang Z, Shi J. Accelerating the discovery of materials with expected thermal conductivity via a synergistic strategy of DFT and interpretable deep learning. Mater. Futures 4:045602 (2025) https://doi.org/10.1088/2752-5724/ae08d0
- E. Atanassov, S. Ivanovska, On the Use of Sobol' Sequence for High Dimensional Simulation, in: D. Groen, C. de Mulatier, M. Paszynski, V.V. Krzhizhanovskaya, J.J. Dongarra, P.M.A. Sloot (Eds.), Computational Science - ICCS 2022, Springer International Publishing, Cham, 2022: pp. 646-652 https://doi.org/10.1007/978-3-031-08760-8_53
- Richardson RR, Osborne MA, Howey DA. Gaussian process regression for forecasting battery state of health. Journal of Power Sources 357:209-219 (2017) https://doi.org/10.1016/j.jpowsour.2017.05.004
- S.W. Ober, C.E. Rasmussen, M. van der Wilk, The Promises and Pitfalls of Deep Kernel Learning, (2021) https://doi.org/10.48550/arXiv.2102.12108
- C.E. Rasmussen, C.K.I. Williams, Gaussian processes for machine learning, 3. print, MIT Press, Cambridge, Mass., 2008.
- E. Brochu, V.M. Cora, N. de Freitas, A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning, (2010) https://doi.org/10.48550/arXiv.1012.2599
- S. Ament, S. Daulton, D. Eriksson, M. Balandat, E. Bakshy, Unexpected Improvements to Expected Improvement for Bayesian Optimization, (2025) https://doi.org/10.48550/arXiv.2310.20708
- X. Jin, K.A. High, Decision Making for a Sustainable Chemical Process, (n.d.).
- T. Akiba, S. Sano, T. Yanase, T. Ohta, M. Koyama, Optuna: A Next-generation Hyperparameter Optimization Framework, (n.d.).
- L. McInnes, J. Healy, J. Melville, UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, (n.d.).
- Ikotun AM, Ezugwu AE, Abualigah L, Abuhaija B, Heming J. K-means clustering algorithms: a comprehensive review, variants analysis, and advances in the era of big data. Information Sciences 622:178-210 (2023) https://doi.org/10.1016/j.ins.2022.11.139
- L.N. Sridhar, Multiobjective Nonlinear Model Predictive Control of the Williams Otto Process, Chemical Engineering 1 (2024).
(0.13 seconds)
[0.14 s]

