Browse
Recent Submissions
New records verified within the last 30 days
Showing records 16 to 40 of 65. [First] Page: 1 2 3 Last
Efficient Parameter Estimation In Agent-Based Models of Collective Cell Invasion Via Gaussian Process Surrogates and Bayesian Optimization
Aneesh Krishna
July 13, 2026 (v1)
Cell migration and invasion are key processes underlying cancer metastasis, driven by cell-cell adhesion, chemotaxis, and matrix remodeling. Agent-based models (ABMs) are a powerful computational approach for simulating these complex, multicellular behaviors. In an ABM, each cell is represented as an autonomous agent that follows local rules governing its interactions with neighboring cells and the surrounding environment. This bottom-up framework captures the emergent, heterogeneous dynamics of cell populations that continuum models cannot easily reproduce. CompuCell3D is a widely used platform for ABMs in biological cell systems; however, a critical limitation is that it and most other ABM tools do not natively support systematic, black-box parameter estimation for their stochastic simulations. Identifying biologically meaningful parameter sets for ABMs is especially challenging because the stochastic nature of these simulations means that repeated runs with identical parameters prod... [more]
From Disparate Data to Accelerating Innovation: A Practical Framework for R&D Digitalization
Zifeng Li
July 13, 2026 (v1)
Manufacturing has benefited from decades of digital standardization, automation, and mature data pipelines. Digitalization in research and development (R&D)-especially in materials and process development- lags significantly due to its heterogeneous and continuously evolving datasets spanning structured measurements, semi‑structured metadata, and unstructured content such as lab notes. The lack of flexible, end‑to‑end digital infrastructure leads to common pain points, including data loss, repeated experiments, long cycle times to derive insight, and barriers to finding and reusing prior knowledge. This work presents a practical digital transformation framework for R&D, centered on knowledge‑centric platforms designed and operated as scalable digital products. At Qnity, an R&D digital transformation project focused on four tightly connected areas.First, customer insight, where direct customer feedback is translated into scalable physics‑based and AI models for product development, supp... [more]
Real-Time Inline Nitric Acid Quantification In Purex Systems Using Raman, ATR-FTIR, and Machine Learning
Nischal Maharjan
July 13, 2026 (v1)
Liquid-liquid extraction (LLE) systems used in spent nuclear fuel reprocessing require reliable, real-time monitoring methods to improve process control, operational efficiency, and material accountancy. Spectroscopic techniques combined with chemometric and machine learning approaches provide a promising pathway for rapid, non-destructive quantification in these complex biphasic environments. In this work, PUREX-relevant solvents, nitric acid solutions (≤ 5 M) and 30% v/v tributyl phosphate (TBP) in n-dodecane, were used as a model LLE system to develop a chemometric workflow for direct quantification of nitric acid extraction from mixed-phase Raman spectra without phase separation. In parallel, single-phase aqueous and organic measurements were collected using both Raman and attenuated total reflectance Fourier transform infrared (ATR-FTIR) spectroscopy to evaluate sensor-specific chemometric performance across relevant concentration ranges. Different machine learning methods were ap... [more]
Tennet-SAC: A Physics-Embedded Machine Learning Model for Activity Coefficient of Multicomponent Liquid Mixtures
Shiang-Tai Lin
July 13, 2026 (v1)
We present TeNNet-SAC (Thermodynamics-embedded Neural Network for Segment Activity Coefficient) [1], a physics-embedded machine learning framework for predicting activity coefficients in multicomponent liquid mixtures directly from SMILES. Building upon the concept of segment-based thermodynamic foundation of COSMO-SAC model [2], TeNNet-SAC preserves physical interpretability while eliminating the need for quantum chemical calculations. The model comprises three components: (i) a σ-profile predictor that infers molecular surface charge distributions from SMILES, (ii) a geometry predictor for molecular volume and surface area, and (iii) a Γ predictor that computes segment activity coefficients. The σ-profile and geometry predictors are trained on 39,745 chemically diverse quantum-calculated structures, ensuring broad chemical coverage. The Γ predictor is designed to enforce thermodynamic consistency and is pretrained on one million synthetic data points to reproduce segment activity coe... [more]
Empowering Automated Process Analysis through LLM-Based Literature Mining, Flowsheet Digitization, and Simulation
Jan-Frederic Laub
July 13, 2026 (v1)
The chemical industry needs to transition from predominantly linear, carbon-emitting production routes to circular, carbon-reusing processes. Therefore, every current and future production process needs to be critically evaluated and potentially re-designed. Today, process design and assessment rely on detailed process simulations [1]. However, constructing these simulations remains a bottleneck, demanding a high degree of expertise and manual work. Here, we present an automated workflow which gathers process knowledge, generates process simulations, and evaluates process performance. The automated workflow consists of two data pipelines: The first pipeline systematically extracts information from literature [2] and prepares a knowledge base of established, industrially relevant chemical processes down to the level of unit operations and thermodynamic properties. This knowledge is aggregated into one text per process and fed into the second pipeline, "text2flowsheet" [3], which digitiz... [more]
Orchestrating Modelling & Simulation of Pharmaceutical Production Processes Via Large Language Models
Christoph Kloss
July 13, 2026 (v1)
The digitalization of production processes necessitates the transition from siloed simulation tools to integrated, automated, and intelligent workflows. In the domain of particulate processes in pharmaceutical manufacturing, agitated drying is playing an important role. Predicting crystal attrition is critical to maintaining target particle size distributions (PSDs) and bioavailability. Traditionally, detailed 3D discrete element method (DEM) models and reduced-order process models have remained separate due to computational disparities. This work proposes an integrated approach to bridge these scales by leveraging physics-informed models and Large Language Models (LLMs) to orchestrate complex characterization and optimization loops. The first time Machine Learning was used to accelerate the characterization of powder properties for DEM was by Benvenuti et al. (2016) . Subsequent innovations introduced physics-informed reduced-order models, such as the two-dimensional population balanc... [more]
How to Teach Programming to ChE's In the Age of AI
Robert Hesketh
July 13, 2026 (v1)
At Rowan University we decided to address a perennial student complaint that the computer science programming class material was never used in later chemical engineering classes. In 2022, we replaced the required programming course with a required chemical engineering course called ChE Modeling. This course introduces students to the modeling of chemical processes using practical simulation tools; the same ones used in industry. Students learn how to build models of complex chemical processes, evaluate the accuracy of models, and use models for process optimization and design decisions. We start this course using the Begin Python with TCLAB[i] modules. This is a unique module in which they learn a programming language to control an Arduino that has 2 heaters and thermistors. This immediately addresses a common complaint that they never used the programming language taught by computer science in a chemical engineering class; they now use python immediately. John Hedengren, the developer... [more]
Why Transfer Learning Fails Under Target Non-Identifiability
Yuki Kobayashi
July 13, 2026 (v1)
Transfer learning (TL) improves model performance in a target domain with limited data by leveraging data from a source domain. When the source-target discrepancy is large, TL can degrade target-domain performance, a phenomenon known as negative transfer (NT). In linear regression, the coefficient vector is only partially identifiable when the target design matrix is rank-deficient. Although TL can exploit source information to address such non-identifiability, it may also amplify coefficient estimation error. However, the mechanisms and conditions underlying this type of NT have not been fully characterized. Frustratingly easy domain adaptation (FEDA) is a TL method that has been successfully applied in the process industry. This study derives the mechanism and conditions of NT in FEDA with linear regression. Because the derived NT condition involves the unobservable true target coefficient, we further construct a proxy condition that can be evaluated from observed data by assuming an... [more]
Domain-Decomposition Pinns for Rapid Prediction of Stirred-Tank Mixing Flows across Geometric Scales
Yohei Kono
July 13, 2026 (v1)
Stirred-tank mixing is a central operation in batch chemical process design and scale-up. While high-fidelity Computational Fluid Dynamics (CFD) is effective for evaluating internal flow states, its computational cost prohibits its use in many-query tasks such as design-space exploration and parametric screening across varied equipment scales. Recent advances in Physics-Informed Neural Networks (PINNs) offer a promising path toward rapid surrogate modeling; however, conventional PINNs often fail to capture complex boundary conditions near impeller regions, frequently collapsing to trivial zero-velocity solutions due to dominant localized forcing terms. In this work, we present a data-informed, physics-constrained hybrid surrogate model for stirred-tank mixing within a bounded geometric-operating space. We employ a domain-decomposition approach: the complex impeller-region dynamics are represented through data-driven boundary conditions derived from a limited set of initial CFD simulati... [more]
A Streamlit-Based Platform for Ternary Solvent Solubility Modeling and Crystallization Process Design
Marko Ivancevic
July 13, 2026 (v1)
Crystallization design in pharmaceutical development is commonly supported by modeling approaches for single‑solvent and binary solvent systems. In some cases, ternary solvent mixtures are desirable due to better impurity purge. However, this benefit can be offset by increased operational and analytical complexity arising from the multicomponent solvent composition, which can limit systematic exploration and broader adoption of ternary crystallization strategies during process development. To address this challenge, a user‑accessible analytics platform has been developed using Streamlit to enable ternary solvent solubility modeling and crystallization process design within an integrated workflow. The application allows users to upload experimentally measured solubility data for ternary solvent systems via a web‑based interface. Upon data import, the software automatically fits the data to five semi‑empirical ternary solubility models using nonlinear regression. Model performance is eva... [more]
Stitching Misoriented and Misaligned Non-Overlapping Images
Michael Fokuo
July 13, 2026 (v1)
Image stitching is a fundamental task in biomedical imaging, enabling reconstruction of large specimens that exceed the field of view of a single acquisition. Conventional stitching methods rely on overlap between neighboring tiles and known acquisition geometry. These methods typically use feature matching registration to estimate spatial transformations and alignment [1-3]. Although effective under controlled conditions, these assumptions do not hold when tiles do not overlap or when their orientations are unknown. This situation often occurs during the physical handling of samples [4]. We address this challenge by proposing a framework for stitching non-overlapping biomedical image tiles with unknown orientations. We conduct the study using a single biomedical dataset that is divided into two subsets: (1) a translation-only (alignment) subset, and (2) a translation-and-orientation subset. In the first subset, orientations are known, but both horizontal and vertical offsets between t... [more]
How Archimetis Operational Reasoning System Detected a Hidden Furnace Failure In Under an Hour
Charles Crowell
July 13, 2026 (v1)
A mid-sized refinery (150K BPD) was experiencing repeated thermal cycles and unit trips on its Gasoline Hydrotreater Reactor Furnace. The plant's DCS and APC indicated the furnace was operating within healthy parameters, masking a serious latent failure. This poster presents how Archimetis's Operational Reasoning System (ORS), a multi-agent system that ingests heterogeneous plant data and reasons over it the way an experienced process engineer would identified the root cause in under an hour. Working with plant engineers, the ORS reviewed several months of operating data alongside event data from the cycles and trips. It detected a 3% gap between theoretical and measured stack O₂, indicating an air-deficient, fuel-rich combustion state inconsistent with the air flow transmitter reading. By cross-referencing combustion stoichiometry, draft pressure, control-valve backpressure trends, and radiant-versus-convective duty shifts, the ORS isolated the cause to a major Air Preheater (APH) lea... [more]
Science Guided Machine Learning for Conceptual Process Development: A Novel Heterogeneous Azeotropic Separation Case Study
Troy Gustke
July 13, 2026 (v1)
Hybrid modeling has stood out as stable method to apply machine learning (ML) to chemical process systems, as it can apply first-principles knowledge and constraints to data-driven methods for increased speed and precision. While substantial work has focused on leveraging operational data, the application of ML to conceptual process design remains limited, despite its strong influence on overall plant economics. Various data-driven methods have improved the speed and accuracy of process development, particularly for optimization; however, they are typically applied to established processes or predefined superstructures, leaving the conceptual process synthesis stage largely unaddressed. In this work, we integrate existing and novel science-guided ML methods into a process development framework to address challenging design problems. This approach integrates with existing process development procedures while benefitting from novel data-driven tools. These methods include modern ML therm... [more]
Role of Multivariate Data Analyses In Formulation and Process Development of Oral Solid Drug Products: Encapsulation Case Studies
Shashwat Gupta
July 13, 2026 (v1)
Multivariate data analysis (MVDA) methods are an important tool in a pharmaceutical drug product engineer's toolbox as part of a Quality by Design (QbD) driven framework for formulation and process development of oral solid dosage forms. This work presents such an MVDA application for two commonly used encapsulation processes in the following case studies: Case study 1: Partial least squares regression (PLSR) based enhancement of process efficiency of a vacuum-assisted drum filling encapsulation process A vacuum-assisted drum filling process was used to fill an active pharmaceutical ingredient (API). Intuitively, the drum bore volume is a function of the amount of API to be filled (combination of dose and incoming potency) and API physical properties (particle size distribution, density, etc.). Thus, selection of an appropriate drum bore volume was desired to reduce process setup time. To enable this, a PLSR model was built correlating API physical properties and drum filling process p... [more]
DEM-CFD Modeling of a Packed Bed Reactor: Analysis of Local Transport & Deactivation Dynamics across Aspect Ratios
Raj Chapagain
July 13, 2026 (v1)
Keywords: AN-SYS Fluent, Catalyst Deactivation, CFD-DEM, Discrete Element Method, Heat Transfer, Low Aspect Ratio, Nusselt Number, Packed Bed Reactor, Porosity, Pressure Drop, Reynolds Stress Anisotropy, Rocky DEM, Thermal Runaway, Turbulence Modelling
Our thesis presents a three-dimensional computational investigation of coupled fluid flow (in an ANSYS 2-Way Fluent Coupling, student version), heat and mass transfer, and catalyst deactivation dynamics within a low aspect ratio packed bed reactor (D/dp < 6), where confining wall effects and packing heterogeneity govern local transport phenomena. The work bridges the gap between molecular-scale reaction kinetics and reactor-scale performance through an integrated Discrete Element Method-Computational Fluid Dynamics (CFD-DEM) modelling framework. The random packing of M = 243 spherical catalyst particles inside a cylindrical vessel (diameter D = 0.15 m, height H = 1.2 m) was generated using DEM with the Hertz-Mindlin contact model and Coulomb friction (μs = 0.5), replicating an industrial gravity-driven settling process. The resulting geometry was transferred via Boolean subtraction to create the interstitial fluid domain, which was discretised using mixed tetrahedral-polyhedral element... [more]
From Mechanistic Model to Digital Twin: A Framework for Real-Time Optimization of Ethanol Production In S.Cerevisiae
Omar Bayomie
July 13, 2026 (v1)
The transition to smart bioprocessing requires control strategies capable of managing the nonlinear dynamics and limited observability in industrial fermentation. In this work, we developed a fed-batch digital twin process for ethanol production by Saccharomyces cerevisiae via combining a mechanistic model with advanced data assimilation to achieve robust Nonlinear Model Predictive Control (NMPC). Recursive Bayesian state estimators have been developed to overcome nonlinearities of the biological models both anaerobic and aerobic, complex metabolic shifts, batch-to-batch parameters variability, and lack of biomass online measurements. Observers (soft sensors) were constructed and benchmarked for computational tractability and statistical accuracy, including Extended (EKF), Ensemble (EnKF), and Particle Filters (PF), They achieved best overall MSE improvements over the mechanistic model; The closed-loop framework showed a robust performance tested by parametric model-plant mismatches, a... [more]
Multivariate PAT Monitoring of Sodium Phosphate Solubility and Crystallization In Alkaline Media
Viviana Cardenas Ocampo
July 13, 2026 (v1)
Crystallization in alkaline process streams is challenging to monitor because solubility, aqueous speciation, hydrate form, and crystal morphology can evolve simultaneously with temperature, and pH. Motivated by phosphate-bearing alkaline waste streams relevant to nuclear waste processing, this work contributes to the development of real-time analytical and modeling tools for crystallization-prone process systems. We develop a data-rich Process Analytical Technology framework that integrates in situ spectroscopy, particle monitoring, temperature and pH measurements with multivariate calibration to quantify phosphate species and track solubility and crystallization behavior in real time. In situ experiments were conducted to characterize sodium phosphate solubility and crystallization under two chemical regimes: strongly alkaline conditions, using 3 molal NaOH, and unadjusted-pH conditions without added NaOH. These conditions provide access to different phosphate speciation regimes. Onl... [more]
Data Driven Experimental Design of Cellulose and Chitin-Based Sustainable Barrier Films
Jessica Bonsu
July 13, 2026 (v1)
The widespread use of non-degradable petroleum-based plastics in food packaging, favored for their cost-effectiveness and strong barrier properties, has led to severe environmental issues. Cellulose and chitin, the most abundant polysaccharides in nature, offer promising sustainable alternatives. Their nanomaterials exhibit high crystallinity and strong hydrogen bonding, which enable excellent mechanical and barrier performance. This makes cellulose- and chitin-based nanomaterials ideal candidates for developing barrier films to substitute petroleum-based plastics. Developing high-performance barrier films requires an understanding of the process-structure-property (PSP) relationships that govern oxygen and moisture transport as well as mechanical stability. This work applies data-driven materials informatics to accelerate that understanding. A curated dataset of 104 cellulose- and chitin-based barrier films, compiled from both literature and laboratory experiments, was constructed to... [more]
A Novel Uncertainty-Aware Computer Vision Framework for Automated Process Optimization In Additive Manufacturing
Ronald Borja-Roman
July 13, 2026 (v1)
The thermomechanical response of materials under high-strain deformation is critical to aerospace, defense, automotive, and additive manufacturing (AM) applications [1-3]. However, capturing this behavior remains challenging due to microsecond timescales of high-velocity impact events and reliance on manual, operator-dependent post-processing [4, 5]. These limitations reduce reproducibility and constrain the generation of high-fidelity datasets for constitutive model development. Current approaches are also limited in accessible strain-rate regimes, restricting applicability to AM processes and introducing uncertainty in predictive modeling [2]. This work presents an autonomous, AI-driven framework that transforms raw high-speed impact videos into structured, model-ready material data. The framework integrates two vision foundation models: Grounding DINO for open-vocabulary object detection and the Segment Anything Model (SAM) for high-resolution segmentation, in a zero-shot configurat... [more]
Leveraging Machine Learning for Multi-Level Optimization In Energy-Water Nexus Systems
Elizabeth Abraham
July 13, 2026 (v1)
The energy-water nexus emerged in response to global challenges associated with energy and water resources. Along with their existing residential, commercial, and industrial commitments, the intrinsically complex system now faces additional pressure from the rampant rise in demands from data centers. To factor in the impacts of these new circumstances while accounting for the interdependent and interconnected nature of energy and water supply systems, the nexus holistically manages these resources and their corresponding decisions [1]. However, while these decisions are conventionally modeled from a centralized perspective through representative mathematical programs whose optimal decisions can then be optimized, a more realistic perspective models these decisions sequentially. Here, rather than all involved systems making their decision in a simultaneous fashion, decisions are made one after another in sequential order and characterized using multilevel programming. Bilevel programmin... [more]
Digital AI-Driven Methodologies to Support and Accelerate Mabs Development In the Biopharmaceutical Industry
Gianmarco Barberi
July 13, 2026 (v1)
Monoclonal antibodies (mAbs) have become a cornerstone in the treatment of immunological and oncological diseases. Nevertheless, bringing new antibody therapeutics to market remains a costly and time-consuming process, often requiring more than 10 years of development and investments exceeding 2 billion dollars. Critical stages of the development pipeline include the identification of a robust production cell line, capable of ensuring key quality attributes such as productivity, stability, product quality, and production consistency, as well as the optimization of the culture process (Li et al., 2010). These steps typically rely on extensive experimental campaigns and significant resource allocation. As a result, pharmaceutical companies are increasingly investigating digital and AI-driven approaches to streamline development workflows and accelerate drug time-to-market. In this work, we address two key challenges in mAb process development: i) the automated detection of anomalous cell... [more]
Expanding the Science-Guided Machine Learning Applications for Process Industries: Advances, Education, and Workforce Development
Y. A. Liu
July 13, 2026 (v1)
The chemical engineering (ChE) field is becoming increasingly hybrid digital as new advancements in machine learning (ML) integrate with ChE workflows and industrial plant operations. As data scientists and engineers push artificial intelligence (AI) usage, it is important for the engineering workforce and data science methods to be grounded in ChE fundamentals through methods like science guided machine learning (SGML). SGML is a broad term that includes ML architectures that are informed/embedded with physics, chemistry, and thermodynamics first-principles. This presentation highlights and demonstrates selective SGML advances from 2022 to 2026, including multicomponent phase equilibria, surrogate modeling, uncertainty quantification, and agentic large language models (LLMs). Furthermore, we discuss how each of these advancements accelerates the process design and development workflow, and we demonstrate solving novel process design and separation problems using this workflow, pushing... [more]
Designing the Future Engineer: How AI Is Transforming Learning, Work, and Discovery
John Kitchin
July 13, 2026 (v1)
Artificial intelligence is reshaping what it means to be an engineer and it is changing how we learn, design, and discover. This talk explores the convergence of generative AI, data-driven modeling, and autonomous experimentation, and what that means for engineering education and workforce development. From intelligent tutors and code-generation assistants to self-driving laboratories and agentic research systems, AI is expanding both the cognitive and creative boundaries of the profession. Drawing on examples from open-source educational ecosystems such as pycse, and recent research on agentic science and generative optimization, we will discuss how future engineers can be trained not only to use AI tools, but to think with them, integrating computation, ethics, and domain expertise into continuous, collaborative learning. We will also discuss challenges of buy-in, resources, resistance from both faculty and students, and the need to maintain a balance of traditional learning approach... [more]
Local-Global Learning of Interpretable Control Polices: The Interface between MPC and Reinforcement Learning
Ali Mesbah
July 13, 2026 (v1)
Optimal decision-making under uncertainty is a shared challenge across modern chemical, manufacturing, and energy systems that increasingly demand safe, data-driven autonomy. This talk revisits optimal control through the lens of the Bellman equation, emphasizing how optimal control theory and reinforcement learning have developed complementary, yet largely disconnected, perspectives on global optimality. In one view, central to reinforcement learning, the Bellman equation defines a global optimality condition that guides iterative policy learning from interacting with the system, but typically yields opaque control laws that are difficult to interpret, and deploy in safety-critical settings. In another view, widely adopted in model predictive control (MPC), the Bellman equation underpins tractable finite-horizon optimizations that deliver interpretable, constraint-aware, and modular local controllers, yet without explicit guarantees on alignment with global optimality. Building on the... [more]
Data-driven optimization: efficient adaptive learning for self-driving laboratories
Nick Sahinidis
July 13, 2026 (v1)
Self-driving laboratories promise to compress materials-discovery timelines from years to weeks by replacing trial-and-error experimentation with closed-loop, algorithm-guided campaigns. Yet, despite the rapid proliferation of robotic and automation hardware, today's autonomous labs rely almost exclusively on Bayesian optimization (BO) to decide what experiment to run next. BO is a sensible approach to low-dimensional optimization problems with smooth response surfaces, but it struggles in precisely the regimes that matter most for real materials campaigns: tight experimental budgets, dozens of process parameters, mixed-integer choices, hard physical constraints, and noisy expensive measurements. In this talk, I will show how moving from BO to partitioning-based algorithms can substantially improve data efficiency, scale gracefully to dozens of process variables, and handle the constraints and noise that characterize realistic experimental campaigns. I will summarize a recently complet... [more]
Showing records 16 to 40 of 65. [First] Page: 1 2 3 Last
(0.06 seconds)

[0.09 s]