Browse
Recent Submissions
New records verified within the last 30 days
Showing records 1 to 25 of 76. [First] Page: 1 2 3 4 Last
Supplementary material for: Privacy-Preserving Coordinated Operation of Multi-Player Industrial Network Using Secure Aggregation
Akshdeep singh Ahluwalia, Zachary Wilson, Jeffery E. Arbogast, Can Li
July 16, 2026 (v1)
Subject: Optimization
Keywords: Dsitributed Optimization, Privacy
This submission corresponds to the supplementary material for article submitted to FOCAPO titled:
"Privacy-Preserving Coordinated Operation of Multi-Player Industrial Network Using Secure Aggregation".
Hybrid Physics-Informed Neural Networks for Thermal Process Identification and Control
Sahar Hemmati, Mohammadreza Babaei, John Hedengren
July 16, 2026 (v1)
Keywords: Heat Transfer, Model Order Reduction, Model Predictive Control, Physics-Informed Neural Networks, Thermal Systems
Physics-Informed Neural Networks (PINNs) offer a promising approach for integrating first-principles modeling with data-driven methods, especially in dynamic thermal systems. This study introduces a hybrid PINN framework for a one-dimensional heating rod governed by heat trans-fer equations. Unlike traditional PINNs that rely on time-dependent automatic differentiation, this approach employs numerical derivatives to bypass gradient saturation and enhance robustness. The proposed model demonstrates accurate extrapolation and generalization with limited train-ing data and is effectively used as a surrogate in a Model Predictive Control (MPC) framework for rod-tip temperature regulation. Additionally, a plan is outlined to apply physics-informed dimen-sionality reduction and model order reduction to improve computational efficiency and enable real-time application. The findings affirm PINNs' potential as control-oriented reduced models for thermal processes.
Supplementary Material for ``Physics Constrained Machine Learning for modeling of Chemical Refineries"
Rahul Golder, Bimol Nath Roy, MMF Faruque Hasan
July 13, 2026 (v1)
Subject: Optimization
Keywords: Physics Constrained Machine Learning, Pooling Problems, Process Modeling, Process Optimization
Supplementary material describing the refinery optimization problem used in the manuscript ``Physics Constrained Machine Learning for modeling of Chemical Refineries"
Optimal Solvent Mixture Screening with Graph Neural Networks
Yipei Zhao
July 13, 2026 (v1)
Solubility is a critical parameter governing the efficacy and stability of agrochemical active ingredients (AIs) formulated in solvent mixtures. Traditionally, the screening of efficient solvent mixtures heavily relies on empirical, trial-and-error experiments. The non-linear behaviours of complex solutes in multi-component solvent systems remain difficult to predict. Furthermore, while predictive machine learning models have shown success in predicting the interaction between multi-component solvents, the application of multi-component system solubility prediction is lacking in the agrochemical field. To address this gap, our study employs a Graph Neural Network (GNN) to model the behaviour of dissolving systems from both intramolecular and intermolecular perspectives. Rather than altering the underlying structural architecture, this work focuses on the domain-specific application and transferability of the model to a novel, high-standard experimental dataset of agrochemical AIs and t... [more]
A Deepsets-Guided Framework for Learning Job Priorities In Single-Machine Scheduling
Daniel Zhu
July 13, 2026 (v1)
Scheduling is a fundamental decision-making problem in chemical engineering as well as numerous other sectors, arising in manufacturing, energy systems, and supply chains. Many scheduling problems are NP-hard [1], meaning that even small problems are computationally hard to solve deterministically. As a result, existing exact and heuristic methods face trade-offs between scalability, solution quality, and generalizability. This work addresses these limitations through a hybrid machine learning-optimization framework for the single-machine total tardiness scheduling problem (SMTTP). We introduce the DeepSets-Guided Scheduling Framework (DGSF), a hybrid methodology that integrates data-driven priority learning with structured optimization for single-machine scheduling. First, we propose a geometric instance classification rule that characterizes scheduling instances through aggregate structural parameters, enabling models trained on small instances to generalize to larger instances withi... [more]
Identifiability of Microkinetic Parameters from Multimodal Operando Data
Gabriel Sabença Gusmão
July 13, 2026 (v1)
Chemical kinetics provides the phenomenological framework for the elucidation of reaction mechanisms, in which ab-initio microkinetic models translate density-functional energetics into catalytic rates. Yet the underlying barriers carry uncertainties of 0.1 to 0.3 eV, and the extent to which a given set of measurements can retrieve them remains largely unquantified. Here, we frame the operando inverse problem as a maximum-likelihood estimation over a differentiable microkinetic model. The pseudo-steady-state surface enters as an algebraic constraint, and parameter sensitivities follow by automatic differentiation through its adjoint, the implicit function theorem applied at the converged root rather than through the solver iterations. These sensitivities propagate the measurement covariance into the parameter covariance, and the resulting Fisher information, the information each experiment carries about each barrier, ranks candidate experiments to establish which measurement determines... [more]
Topology-Guided Response Surface Characterization for ML Model Selection
Shenbageshwaran Rajendiran
July 13, 2026 (v1)
Machine learning (ML) models are widely used as surrogates for complex optimization problems in process systems engineering; however, selecting an appropriate surrogate remains challenging. Current selection approaches often rely on cross-validation, prior experience, or statistical and gradient-based response-surface descriptors. We hypothesize that topological descriptors such as the Euler characteristic curve (ECC), which captures structural changes across thresholds, provide complementary information for structure-aware surrogate model selection. We evaluated this hypothesis using 41 two-dimensional optimization test functions annotated with four landscape characteristics: modality, ruggedness, abruptness, and geometric complexity. For each function, ECCs were first computed on a dense 25*25 (625 samples) grid using sublevel filtration. Summary metrics from ECC, including the number of ECC jumps, cumulative ECC variation, and maximum ECC jump, were extracted from each curve. Thresh... [more]
A Machine Learning Framework for Short Peptide Sequence Optimization
Anh Trinh
July 13, 2026 (v1)
Designing peptides plays an important role in applications ranging from therapeuticsand biomaterials to diagnostics. However, due to the large combinatorial sequence spaceand the high cost and time required for experimental screening, experimental trial anderror approaches are prohibitively expensive. Furthermore, peptide design is inherentlya multi-objective problem that requires simultaneous optimization of different propertiessuch as biological activity, stability, solubility, and safety. These challenges motivate theuse of computational design strategies. Traditional physics-based and sequence-alignmentmethods often struggle to handle variable length sequences and often rely on structuralinformation that is unavailable for many peptides.1, 2 More recently, deep learning modelssuch as AlphaFold, ESM, and diffusion-based approaches have transformed protein mod-eling3, 4, 5 . However, their large data requirements, high computational cost, and black-boxnature reduce their practicality... [more]
Generalized Physics-Informed Deep Learning Framework for Chemical Process Modeling
Harshit Verma
July 13, 2026 (v1)
The incorporation of mechanistic, first-principles chemical unit operation models into process modeling frameworks remains computationally challenging. Mechanistic models governed by complex nonlinear systems of ordinary and partial differential equations are intractable for modern deterministic global solvers, particularly within large-scale nonlinear and mixed-integer nonlinear programming (MINLP) formulations. As a result, surrogate modeling approaches have gained increasing attention. However, conventional surrogate models typically rely on strong process-level assumptions, simplified physics, or extensive data generation, which limits their extrapolation capability, physical consistency, and reliability across feasible process operating regions. Physics-informed neural networks (PINNs) offer a promising alternative by embedding governing physics directly into the learning objective loss function, thereby reducing dependence on large supervised training datasets while preserving go... [more]
Safety System Complexity: Ontology-Grounded Llms That Cross-Link Plant Records to Reduce Spurious Trips and Surface Hidden Process Risk
David Parham
July 13, 2026 (v1)
Instrumented safety systems have driven incident rates down, but also create new challenges: the plants they protect are now far more complex. Two problems follow. First, true process risk hides in combinations of latent failures, process variability, corrosion, and degradation, so catastrophic incidents persist despite layered protection. Second, spurious safety-system trips have risen, eroding availability and driving unplanned downtime. The challenge is to keep the safety benefits while giving operations the information needed to run both safely and reliably. We present a case study in which large language models, grounded in a domain-specific process-safety ontology, extract and cross-link evidence across P&IDs, asset registers, PHAs, management-of-change records, and incident reports. The system surfaces discrepancies between sources and answers operators' questions on demand, at the point of decision, with each answer traceable to its evidence. We report how ontology grounding ch... [more]
Machine Learning-Based Prediction of Heavy Metal Exposure In Spatially Heterogeneous Urban Environments
Paromita Nath
July 13, 2026 (v1)
Machine learning (ML) is increasingly used for predictive analysis in complex environmental systems. However, its effectiveness is often constrained by spatial heterogeneity and resource limitations that restrict data availability. This study evaluates the feasibility and limitations of ML-based predictive modeling for estimating stormwater heavy metal concentrations and associated public health risks in Camden, New Jersey, a historically industrial city with known contamination challenges. In this work, limited stormwater sampling data are integrated with environmental and anthropogenic predictors, including land use, proximity to industries, vegetation, and elevation for machine learning model development. Multiple regression-based ML models, including linear, ridge, lasso, random forest, and support vector regression, are trained and evaluated. In addition, hierarchical modeling approaches are explored to improve predictive performance. Results highlight challenges in predictive gen... [more]
Renewable-Driven Microgrid Design, Planning, and Operation of Integrated Gasification Fuel Cell for Biomass Upgradation to Biofuels
Oluwatimileyin Ogunsola
July 13, 2026 (v1)
Biofuel production and upgrading offer a promising pathway for reducing the carbon intensity of liquid fuels. However, biomass-to-biofuel systems are energy intensive and require coordinated supplies of electricity, heat, and hydrogen. In a typical process, biomass is converted to bio-oil through pyrolysis, hydrogen is produced using an electrolyzer, and the resulting bio-oil is upgraded to transportation-grade biofuel. When the pyrolyzer is electrically heated and the upgrading unit uses an electrochemical pathway, the overall process imposes a large and time-varying electrical demand. Supplying this demand with renewable energy is attractive, but the variability of solar and wind generation can lead to renewable curtailment, grid dependence, and operational challenges. These issues motivate the development of integrated microgrid scheduling frameworks that can coordinate renewable generation, storage, grid exchange, and onsite fuel-based power generation. This work develops a renewab... [more]
Sketch2Simulation: Automating Flowsheet Generation Via Multi-Agent Large Language Models
Emma Pajak
July 13, 2026 (v1)
Converting process flow diagrams into complete simulation models remains a persistent bottleneck in process systems engineering (PSE), requiring significant manual effort and simulator-specific expertise. Although advances in diagram interpretation and automated model generation have been made, these tasks are typically addressed in isolation, limiting the automation of end-to-end workflows. This work introduces Sketch2Simulation, a unified computational framework that automates flowsheet generation directly from raw engineering diagrams using a multi-agent large language model (LLM) architecture. The proposed framework integrates three coordinated layers: (i) Diagram Parsing and Interpretation, (ii) Simulation Model Synthesis, and (iii) Multi-level Validation. In the first layer, multimodal LLM agents extract process semantics, identify unit operations and stream connectivity, and resolve implicit structural features. This information is encoded into a directed graph-based intermediat... [more]
Reinforcement Learning for Nonlinear Optimization In Process Industry
Kalpesh Patel
July 13, 2026 (v1)
Reinforcement Learning (RL) is a machine learning technique which is capable of generating data and learning from it autonomously by interacting with the environment. RL has been successfully applied for learning and playing various games such as Go, Chess, Atari etc but its application to address process control and optimization problems is not trivial. There is a need for RL implementations in process industry to be safe, fast learning and explainable. A method for achieving such an implementation, for linear systems without disturbance variables, was published by the author in the past. Taking the work further, this paper proposes significant enhancements to the method in terms of ability to address severe process non-linearities, that can't be linearized, and ability to address disturbance variables explicitly in the RL problem formulation. Along with presenting the details on the enhancements, the paper also provides details on actual implementation of the enhanced method for opti... [more]
Optimization of Biogas Steam Reforming Toward Low Carbon Hydrogen Production Using Integrated Artificial Neural Network and Genetic Algorithm
Ikechukwu Okwuosa
July 13, 2026 (v1)
Hydrogen has been identified as a versatile energy carrier, offering a viable route to decarbonize and meet escalating global energy demands. Biogas produced from the anaerobic digestion of organic matter can potentially serve as a feedstock for hydrogen production using steam reforming process. This research investigates the optimization of a steam reforming process utilizing biogas feedstock for low-carbon hydrogen production using Artificial Neural Network (ANN) integrated with Genetic Algorithm (GA). An equilibrium based steady-state simulation of the process was developed using Aspen HYSYS to generate data for neural network training, validation and testing. Key process parameters considered for optimization include: biogas flow rate, steam flow rate, reformer temperature and reformer pressure with hydrogen mole fraction at reformer outlet as the response variable. A two-layer feedforward neural network with 4-12-1 architecture was trained on simulation data, achieving a correlati... [more]
Efficient Parameter Estimation In Agent-Based Models of Collective Cell Invasion Via Gaussian Process Surrogates and Bayesian Optimization
Aneesh Krishna
July 13, 2026 (v1)
Cell migration and invasion are key processes underlying cancer metastasis, driven by cell-cell adhesion, chemotaxis, and matrix remodeling. Agent-based models (ABMs) are a powerful computational approach for simulating these complex, multicellular behaviors. In an ABM, each cell is represented as an autonomous agent that follows local rules governing its interactions with neighboring cells and the surrounding environment. This bottom-up framework captures the emergent, heterogeneous dynamics of cell populations that continuum models cannot easily reproduce. CompuCell3D is a widely used platform for ABMs in biological cell systems; however, a critical limitation is that it and most other ABM tools do not natively support systematic, black-box parameter estimation for their stochastic simulations. Identifying biologically meaningful parameter sets for ABMs is especially challenging because the stochastic nature of these simulations means that repeated runs with identical parameters prod... [more]
From Disparate Data to Accelerating Innovation: A Practical Framework for R&D Digitalization
Zifeng Li
July 13, 2026 (v1)
Manufacturing has benefited from decades of digital standardization, automation, and mature data pipelines. Digitalization in research and development (R&D)-especially in materials and process development- lags significantly due to its heterogeneous and continuously evolving datasets spanning structured measurements, semi‑structured metadata, and unstructured content such as lab notes. The lack of flexible, end‑to‑end digital infrastructure leads to common pain points, including data loss, repeated experiments, long cycle times to derive insight, and barriers to finding and reusing prior knowledge. This work presents a practical digital transformation framework for R&D, centered on knowledge‑centric platforms designed and operated as scalable digital products. At Qnity, an R&D digital transformation project focused on four tightly connected areas.First, customer insight, where direct customer feedback is translated into scalable physics‑based and AI models for product development, supp... [more]
Real-Time Inline Nitric Acid Quantification In Purex Systems Using Raman, ATR-FTIR, and Machine Learning
Nischal Maharjan
July 13, 2026 (v1)
Liquid-liquid extraction (LLE) systems used in spent nuclear fuel reprocessing require reliable, real-time monitoring methods to improve process control, operational efficiency, and material accountancy. Spectroscopic techniques combined with chemometric and machine learning approaches provide a promising pathway for rapid, non-destructive quantification in these complex biphasic environments. In this work, PUREX-relevant solvents, nitric acid solutions (≤ 5 M) and 30% v/v tributyl phosphate (TBP) in n-dodecane, were used as a model LLE system to develop a chemometric workflow for direct quantification of nitric acid extraction from mixed-phase Raman spectra without phase separation. In parallel, single-phase aqueous and organic measurements were collected using both Raman and attenuated total reflectance Fourier transform infrared (ATR-FTIR) spectroscopy to evaluate sensor-specific chemometric performance across relevant concentration ranges. Different machine learning methods were ap... [more]
Tennet-SAC: A Physics-Embedded Machine Learning Model for Activity Coefficient of Multicomponent Liquid Mixtures
Shiang-Tai Lin
July 13, 2026 (v1)
We present TeNNet-SAC (Thermodynamics-embedded Neural Network for Segment Activity Coefficient) [1], a physics-embedded machine learning framework for predicting activity coefficients in multicomponent liquid mixtures directly from SMILES. Building upon the concept of segment-based thermodynamic foundation of COSMO-SAC model [2], TeNNet-SAC preserves physical interpretability while eliminating the need for quantum chemical calculations. The model comprises three components: (i) a σ-profile predictor that infers molecular surface charge distributions from SMILES, (ii) a geometry predictor for molecular volume and surface area, and (iii) a Γ predictor that computes segment activity coefficients. The σ-profile and geometry predictors are trained on 39,745 chemically diverse quantum-calculated structures, ensuring broad chemical coverage. The Γ predictor is designed to enforce thermodynamic consistency and is pretrained on one million synthetic data points to reproduce segment activity coe... [more]
Empowering Automated Process Analysis through LLM-Based Literature Mining, Flowsheet Digitization, and Simulation
Jan-Frederic Laub
July 13, 2026 (v1)
The chemical industry needs to transition from predominantly linear, carbon-emitting production routes to circular, carbon-reusing processes. Therefore, every current and future production process needs to be critically evaluated and potentially re-designed. Today, process design and assessment rely on detailed process simulations [1]. However, constructing these simulations remains a bottleneck, demanding a high degree of expertise and manual work. Here, we present an automated workflow which gathers process knowledge, generates process simulations, and evaluates process performance. The automated workflow consists of two data pipelines: The first pipeline systematically extracts information from literature [2] and prepares a knowledge base of established, industrially relevant chemical processes down to the level of unit operations and thermodynamic properties. This knowledge is aggregated into one text per process and fed into the second pipeline, "text2flowsheet" [3], which digitiz... [more]
Orchestrating Modelling & Simulation of Pharmaceutical Production Processes Via Large Language Models
Christoph Kloss
July 13, 2026 (v1)
The digitalization of production processes necessitates the transition from siloed simulation tools to integrated, automated, and intelligent workflows. In the domain of particulate processes in pharmaceutical manufacturing, agitated drying is playing an important role. Predicting crystal attrition is critical to maintaining target particle size distributions (PSDs) and bioavailability. Traditionally, detailed 3D discrete element method (DEM) models and reduced-order process models have remained separate due to computational disparities. This work proposes an integrated approach to bridge these scales by leveraging physics-informed models and Large Language Models (LLMs) to orchestrate complex characterization and optimization loops. The first time Machine Learning was used to accelerate the characterization of powder properties for DEM was by Benvenuti et al. (2016) . Subsequent innovations introduced physics-informed reduced-order models, such as the two-dimensional population balanc... [more]
How to Teach Programming to ChE's In the Age of AI
Robert Hesketh
July 13, 2026 (v1)
At Rowan University we decided to address a perennial student complaint that the computer science programming class material was never used in later chemical engineering classes. In 2022, we replaced the required programming course with a required chemical engineering course called ChE Modeling. This course introduces students to the modeling of chemical processes using practical simulation tools; the same ones used in industry. Students learn how to build models of complex chemical processes, evaluate the accuracy of models, and use models for process optimization and design decisions. We start this course using the Begin Python with TCLAB[i] modules. This is a unique module in which they learn a programming language to control an Arduino that has 2 heaters and thermistors. This immediately addresses a common complaint that they never used the programming language taught by computer science in a chemical engineering class; they now use python immediately. John Hedengren, the developer... [more]
Why Transfer Learning Fails Under Target Non-Identifiability
Yuki Kobayashi
July 13, 2026 (v1)
Transfer learning (TL) improves model performance in a target domain with limited data by leveraging data from a source domain. When the source-target discrepancy is large, TL can degrade target-domain performance, a phenomenon known as negative transfer (NT). In linear regression, the coefficient vector is only partially identifiable when the target design matrix is rank-deficient. Although TL can exploit source information to address such non-identifiability, it may also amplify coefficient estimation error. However, the mechanisms and conditions underlying this type of NT have not been fully characterized. Frustratingly easy domain adaptation (FEDA) is a TL method that has been successfully applied in the process industry. This study derives the mechanism and conditions of NT in FEDA with linear regression. Because the derived NT condition involves the unobservable true target coefficient, we further construct a proxy condition that can be evaluated from observed data by assuming an... [more]
Domain-Decomposition Pinns for Rapid Prediction of Stirred-Tank Mixing Flows across Geometric Scales
Yohei Kono
July 13, 2026 (v1)
Stirred-tank mixing is a central operation in batch chemical process design and scale-up. While high-fidelity Computational Fluid Dynamics (CFD) is effective for evaluating internal flow states, its computational cost prohibits its use in many-query tasks such as design-space exploration and parametric screening across varied equipment scales. Recent advances in Physics-Informed Neural Networks (PINNs) offer a promising path toward rapid surrogate modeling; however, conventional PINNs often fail to capture complex boundary conditions near impeller regions, frequently collapsing to trivial zero-velocity solutions due to dominant localized forcing terms. In this work, we present a data-informed, physics-constrained hybrid surrogate model for stirred-tank mixing within a bounded geometric-operating space. We employ a domain-decomposition approach: the complex impeller-region dynamics are represented through data-driven boundary conditions derived from a limited set of initial CFD simulati... [more]
A Streamlit-Based Platform for Ternary Solvent Solubility Modeling and Crystallization Process Design
Marko Ivancevic
July 13, 2026 (v1)
Crystallization design in pharmaceutical development is commonly supported by modeling approaches for single‑solvent and binary solvent systems. In some cases, ternary solvent mixtures are desirable due to better impurity purge. However, this benefit can be offset by increased operational and analytical complexity arising from the multicomponent solvent composition, which can limit systematic exploration and broader adoption of ternary crystallization strategies during process development. To address this challenge, a user‑accessible analytics platform has been developed using Streamlit to enable ternary solvent solubility modeling and crystallization process design within an integrated workflow. The application allows users to upload experimentally measured solubility data for ternary solvent systems via a web‑based interface. Upon data import, the software automatically fits the data to five semi‑empirical ternary solubility models using nonlinear regression. Model performance is eva... [more]
Showing records 1 to 25 of 76. [First] Page: 1 2 3 4 Last
(0.06 seconds)

[0.08 s]