LAPSE:2026.1228
Published Article
LAPSE:2026.1228
A Machine Learning Framework for Short Peptide Sequence Optimization
Anh Trinh
July 13, 2026
Abstract
Designing peptides plays an important role in applications ranging from therapeuticsand biomaterials to diagnostics. However, due to the large combinatorial sequence spaceand the high cost and time required for experimental screening, experimental trial anderror approaches are prohibitively expensive. Furthermore, peptide design is inherentlya multi-objective problem that requires simultaneous optimization of different propertiessuch as biological activity, stability, solubility, and safety. These challenges motivate theuse of computational design strategies. Traditional physics-based and sequence-alignmentmethods often struggle to handle variable length sequences and often rely on structuralinformation that is unavailable for many peptides.1, 2 More recently, deep learning modelssuch as AlphaFold, ESM, and diffusion-based approaches have transformed protein mod-eling3, 4, 5 . However, their large data requirements, high computational cost, and black-boxnature reduce their practicality for deterministic multi-objective optimization in limited-data settings.6This work proposes a data-driven, multi-objective peptide design framework that inte-grates sequence-to-feature transformations using Fast Fourier Transform (FFT) - basedrepresentations,7, 8 interpretable feature attribution through GroupSHAPLEY, and metric-learning based optimization strategies. A bidirectional mapping between sequence spaceand feature space is introduced to identify critical feature contributions and improve inter-pretability during peptide optimization.The primary focus of this poster is the optimization component of the framework. Specif-ically, distance metric learning methods, including Neighborhood Component Analysis(NCA)9 and Large Margin Nearest Neighbor (LMNN),10 are investigated to maximizeclass separation between peptide groups and identify discriminative feature representations.Comparative analyses were performed to evaluate the robustness, tunability, and optimiza-tion behavior of these methods on both synthetic and peptide-representative datasets. UsingSupport Vector Machine (SVM) as the predictive model, the proposed methods demonstrateimproved optimization efficiency through dimensionality reduction in synthetic data exper-iments. Multiple optimization constraints can be incorporated to assess the robustness,scalability, and adaptability of the framework across different design scenarios. The pro-posed framework aims to provide an interpretable and computationally efficient alternativefor peptide design under limited-data constraints.
Suggested Citation
Trinh A. A Machine Learning Framework for Short Peptide Sequence Optimization. (2026). LAPSE:2026.1228
Author Affiliations
Trinh A: Georgia Institute of Technology, School of Chemical and Biomolecular Engineering
Journal Name
Proceedings of FOPAM 2026
Volume
0
First Page
56
Last Page
56
Year
2026
Publication Date
2026-07-13
Version Comments
Original Submission
Other Meta
PII: 0056-0056-33-PSE-0-2026, Publication Type: Abstract
Record Map
Published Article

LAPSE:2026.1228
This Record
External Link

https://doi.org/10.69997/pse.130642
Publisher Version
Download
Files
Jul 13, 2026
Main Article
License
CC BY-SA 4.0
Meta
Record Statistics
Record Views
153
Version History
[v1] (Original Submission)
Jul 13, 2026
 
Verified by curator on
Jul 13, 2026
This Version Number
v1
Citations
Most Recent
This Version
URL Here
https://psecommunity.org/LAPSE:2026.1228
 
Record Owner
PSE Press
Links to Related Works
Directly Related to This Work
Publisher Version
(0.1 seconds)

[0.1 s]