Finetuning large language model-based protein generator requires only hundreds of protein sequences.
Diversity of finetuning set affects diversity of generated candidates.
Prompting is not necessary to guide the generation.
Coupling a discriminator with protein generator enables efficient design of functional and novel proteins.
We designed experimentally valid antimicrobial peptides and malate dehydrogenases using this framework.
| [1] | Linsky T. W., Vergara R., Codina N., et al. (2020). De novo design of potent and resilient hACE2 decoys to neutralize SARS-CoV-2. Science 370:1208−1214. DOI:10.1126/science.abe0075 |
| [2] | Sesterhenn F., Yang C., Bonet J., et al. (2020). De novo protein design enables the precise induction of RSV-neutralizing antibodies. Science 368:eaay5051. DOI:10.1126/science.aay5051 |
| [3] | Pan X., and Kortemme T. (2021). Recent advances in de novo protein design: Principles, methods, and applications. J. Biol. Chem. 296:100558. DOI:10.1016/j.jbc.2021.100558 |
| [4] | Quijano-Rubio A., Yeh H.-W., Park J., et al. (2021). De novo design of modular and tunable protein biosensors. Nature 591:482−487. DOI:10.1038/s41586-021-03258-z |
| [5] | Liu K., Zhang Y., Liu K., et al. (2022). De novo design of a transcription factor for a progesterone biosensor. Biosens. Bioelectron. 203:113897. DOI:10.1016/j.bios.2021.113897 |
| [6] | Jackson C., Anderson A., and Alexandrov K. (2022). The present and the future of protein biosensor engineering. Curr. Opin. Struct. Biol. 75:102424. DOI:10.1016/j.sbi.2022.102424 |
| [7] | Madhavan A., Arun K., Binod P., et al. (2021). Design of novel enzyme biocatalysts for industrial bioprocess: Harnessing the power of protein engineering, high throughput screening and synthetic biology. Bioresour Technol. 325:124617. DOI:10.1016/j.biortech.2020.124617 |
| [8] | Ennist N. M., Zhao Z., Stayrook S. E., et al. (2022). De novo protein design of photochemical reaction centers. Nat. Commun. 13:4937. DOI:10.1038/s41467-022-32710-5 |
| [9] | Currin A., Swainston N., Day P. J., et al. (2015). Synthetic biology for the directed evolution of protein biocatalysts: navigating sequence space intelligently. Chem. Soc. Rev. 44:1172−1239. DOI:10.1039/c4cs00351a |
| [10] | Ferruz N. and Höcker B. (2022). Controllable protein design with language models. Nat. Mach. Intell. 4:521−532. DOI:10.1038/s42256-022-00499-z |
| [11] | Ferruz N., Heinzinger M., Akdel M., et al. (2022). From sequence to function through structure: deep learning for protein design. Comput. Struct. Biotechnol. J. 21:238−250. DOI:10.1016/j.csbj.2022.11.014 |
| [12] | Madani A., Krause B., Greene E. R., et al. (2023). Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 41:1−8. DOI:10.1038/s41587-022-01618-2 |
| [13] | Min B., Ross H., Sulem E., et al. (2021). Recent advances in natural language processing via large pre-trained language models: A survey. ACM Computing Surveys 56:1−40. DOI:10.1145/3605943 |
| [14] | Budzianowski P. and Vulić I. (2019). Hello, it's GPT-2--how can I help you? towards the use of pretrained language models for task-oriented dialogue systems. arXiv preprint arXiv:1907.05774. DOI:10.48550/arXiv.1907.05774 |
| [15] | Alexandr N., Irina O., Tatyana K., et al. (2021). Fine-tuning GPT-3 for Russian text summarization. Silhavy, R., Silhavy, P., Prokopova, Z. (eds). Data Science and Intelligent Systems. (Springer, Cham), pp:748–757. DOI:10.1007/978-3-030-90321-3_61 |
| [16] | Liao Y., Wang Y., Liu Q., et al. (2019). Gpt-based generation for classical chinese poetry. arXiv preprint arXiv:1907.00151. DOI:10.48550/arXiv.1907.00151. |
| [17] | Strokach A. and Kim P. M. (2022). Deep generative modeling for protein design. Curr. Opin. Struct. Biol. 72:226−236. DOI:10.1016/j.sbi.2021.11.008 |
| [18] | Wang Y., Deng J., Sun A., et al. (2022). Perplexity from PLM Is Unreliable for Evaluating Text Quality. arXiv preprint arXiv:2210.05892. DOI:10.48550/arXiv.2210.05892. |
| [19] | Newman D., Noh Y., Talley E., et al. (2010). Evaluating topic models for digital libraries. Proceedings of the 10th annual joint conference on Digital libraries, pp:215-224. DOI:10.1145/1816123.1816156 |
| [20] | Chang J., Gerrish S., Wang C., et al. (2009). Reading tea leaves: How humans interpret topic models. Advances in neural information processing systems, pp:288-296. |
| [21] | Repecka D., Jauniskis V., Karpus L., et al. (2021). Expanding functional protein sequence spaces using generative adversarial networks. Nat. Mach. Intell. 3:324−333. DOI:10.1038/s42256-021-00310-5 |
| [22] | Dauparas J., Anishchenko I., Bennett N., et al. (2022). Robust deep learning–based protein sequence design using ProteinMPNN. Science 378:49−56. DOI:10.1126/science.add2187 |
| [23] | Yu T., Cui H., Li J. C., et al. (2023). Enzyme function prediction using contrastive learning. Science 379:1358−1363. DOI:10.1126/science.adf2465 |
| [24] | Nivón L. G., Moretti R., and Baker D. (2013). A Pareto-optimal refinement method for protein design scaffolds. PloS one 8:e59004. DOI:10.1371/journal.pone.0059004 |
| [25] | Meier J., Rao R., Verkuil R., et al. (2021). Language models enable zero-shot prediction of the effects of mutations on protein function. Advances in Neural Information Processing Systems 34:29287−29303. |
| [26] | Yang K. K., Fusi N., and Lu A. X. (2022). Convolutions are competitive with transformers for protein sequence pretraining. Cell Syst. 15:286−294.e2. DOI:10.1016/j.cels.2024.01.008 |
| [27] | Yang K. K., Zanichelli N., and Yeh H. (2022). Masked inverse folding with sequence transfer for protein representation learning. Protein Eng. Des. Sel. 36:gzad015. DOI:10.1093/protein/gzad015 |
| [28] | Hsu C., Verkuil R., Liu J., et al. (2022). Learning inverse folding from millions of predicted structures. International Conference on Machine Learning. Proceedings of the 39th International Conference on Machine Learning 162:8946-8970. |
| [29] | Johnson S. R., Fu X., Viknander S., et al. (2025). Computational scoring and experimental evaluation of enzymes generated by neural networks. Nat. Biotechnol. 43:396−405. DOI:10.1038/s41587-024-02214-2 |
| [30] | Villegas-Morcillo A., Makrodimitris S., van Ham R. C., et al. (2021). Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function. Bioinformatics 37:162−170. DOI:10.1093/bioinformatics/btaa701 |
| [31] | Heinzinger M., Elnaggar A., Wang Y., et al. (2019). Modeling aspects of the language of life through transfer-learning protein sequences. BMC bioinformatics 20:1−17. DOI:10.1186/s12859-019-3220-8 |
| [32] | Jha K., Saha S., and Singh H. (2022). Prediction of protein–protein interaction using graph neural networks. Sci. Rep. 12:8360. DOI:10.1038/s41598-022-12201-9 |
| [33] | Singh S., Chaudhary K., Dhanda S. K., et al. (2016). SATPdb: a database of structurally annotated therapeutic peptides. Nucleic Acids Res. 44:D1119−D1126. DOI:10.1093/nar/gkv1114 |
| [34] | Consortium T. U. (2021). UniProt: the universal protein knowledgebase in 2021. Nucleic Acids Res. 49:D480−D489. DOI:10.1093/nar/gkaa1100 |
| [35] | Sidorczuk K., Gagat P., Pietluch F., et al. (2022). Benchmarks in antimicrobial peptide prediction are biased due to the selection of negative data. Briefings Bioinform. 23:bbac343. DOI:10.1093/bib/bbac343 |
| [36] | McDonald A. G., and Tipton K. F. (2021). Enzyme nomenclature and classification: The state of the art. FEBS J. 290:2214−2231. DOI:10.1111/febs.16274 |
| [37] | Li W. and Godzik A. (2006). Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics 22:1658−1659. DOI:10.1093/bioinformatics/btl158 |
| [38] | Wolf T., Debut L., Sanh V., et al. (2019). Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771. DOI:10.48550/arXiv.1910.03771 |
| [39] | Elnaggar A., Heinzinger M., Dallago C., et al. (2020). ProtTrans: towards cracking the language of Life's code through self-supervised deep learning and high performance computing. arXiv preprint arXiv:2007.06225. DOI:10.48550/arXiv.2007.06225 |
| [40] | Chollet F. (2015). Keras: Deep learning library for theano and tensorflow. https://keras.io |
| [41] | Agarap A. F. (2018). Deep learning using rectified linear units (relu). arXiv preprint arXiv:1803.08375. DOI:10.48550/arXiv.1803.08375 |
| [42] | Rasamoelina A. D., Adjailia F., and Sinčák P. (2020). A review of activation function for artificial neural network. 2020 IEEE 18th World Symposium on Applied Machine Intelligence and Informatics (SAMI). IEEE. DOI:0.1109/SAMI48414.2020.9108717. |
| [43] | Kingma D. P., and Ba J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980. DOI:10.48550/arXiv.1412.6980 |
| [44] | Mistry J., Chuguransky S., Williams L., et al. (2021). Pfam: The protein families database in 2021. Nucleic Acids Res. 49:D412−D419. DOI:10.1093/nar/gkaa913 |
| [45] | Eddy S. R. (2011). Accelerated profile HMM searches. PLoS Comput. Biol. 7:e1002195. DOI:10.1371/journal.pcbi.1002195 |
| [46] | Camacho C., Coulouris G., Avagyan V., et al. (2009). BLAST+: Architecture and applications. BMC Bioinformatics 10:1−9. DOI:10.1186/1471-2105-10-421 |
| [47] | Mahlich Y., Steinegger M., Rost B., et al. (2018). HFSP: high speed homology-driven function annotation of proteins. Bioinformatics 34:i304−i312. DOI:10.1093/bioinformatics/bty262 |
| [48] | Hinton G. E. and Roweis S. (2002). Stochastic neighbor embedding. Advances in neural information processing systems, pp:857 - 864. |
| [49] | Pedregosa F., Varoquaux G., Gramfort A., et al. (2011). Scikit-learn: Machine learning in Python. the Journal of machine Learning research 12:2825−2830. |
| [50] | Burley S. K., Bhikadiya C., Bi C., et al. (2023). RCSB Protein Data Bank (RCSB. org): delivery of experimentally-determined PDB structures alongside one million computed structure models of proteins from artificial intelligence/machine learning. Nucleic Acids Res. 51:D488−D508. DOI:10.1093/nar/gkac1077 |
| [51] | Katoh K., Rozewicki J. and Yamada K. D. (2019). MAFFT online service: Multiple sequence alignment, interactive sequence choice and visualization. Brief. Bioinform. 20:1160−1166. DOI:10.1093/bib/bbx108 |
| [52] | Bodenhofer U., Bonatesta E., Horejš-Kainrath C., et al. (2015). msa: An R package for multiple sequence alignment. Bioinformatics 31:3997−3999. DOI:10.1093/bioinformatics/btv494 |
| [53] | Yuan S., Chan H. S., and Hu Z. (2017). Using PyMOL as a platform for computational drug design. Wiley Interdiscip. Rev.:Comput. Mol. Sci. 7:e1298. DOI:10.1002/wcms.1298 |
| [54] | Koh H. Y., Nguyen A. T., Pan S., et al. (2024). Physicochemical graph neural network for learning protein–ligand interaction fingerprints from sequence data. Nat. Mach. Intell. 6:673−687. DOI:10.1038/s42256-024-00847-1 |
| [55] | Koes D. R., Baumgartner M. P., and Camacho C. J. (2013). Lessons learned in empirical scoring with smina from the CSAR 2011 benchmarking exercise. J. Chem. Inf. Model. 53:1893−1904. DOI:10.1021/ci300604z |
| [56] | Lin Z., Akin H., Rao R., et al. (2023). Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379:1123−1130. DOI:10.1126/science.ade2574 |
| [57] | Robin X., Turck N., Hainard A., et al. (2011). pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics 12:1−8. DOI:10.1186/1471-2105-12-77 |
| [58] | Xu J., Li F., Leier A., et al. (2021). Comprehensive assessment of machine learning-based methods for predicting antimicrobial peptides. Brief. Bioinform. 22:bbab083. DOI:10.1093/bib/bbab083 |
| [59] | Devlin J., Chang M.-W., Lee K., et al. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. DOI:10.48550/arXiv.1810.04805 |
| [60] | Palermo E. F., and Kuroda K. (2010). Structural determinants of antimicrobial activity in polymers which mimic host defense peptides. Appl. Microbiol. Biotechnol. 87:1605−1615. DOI:10.1007/s00253-010-2687-z |
| [61] | Shagaghi N., Palombo E. A., Clayton A. H., et al. (2018). Antimicrobial peptides: biochemical determinants of activity and biophysical techniques of elucidating their functionality. World J. Microbiol. Biotechnol. 34:1−13. DOI:10.1007/s11274-018-2444-5 |
| [62] | Kumar P., Kizhakkedathu J. N., and Straus S. K. (2018). Antimicrobial peptides: diversity, mechanism of action and strategies to improve the activity and biocompatibility in vivo. Biomolecules 8:4. DOI:10.3390/biom8010004 |
| [63] | Ma Y., Guo Z., Xia B., et al. (2022). Identification of antimicrobial peptides from the human gut microbiome using deep learning. Nat. Biotechnol. 40:921−931. DOI:10.1038/s41587-022-01226-0 |
| [64] | Nagarajan D., Nagarajan T., Roy N., et al. (2018). Computational antimicrobial peptide design and evaluation against multidrug-resistant clinical isolates of bacteria. J. Biol. Chem. 293:3492−3509. DOI:10.1074/jbc.M117.805499 |
| [65] | Rosano G. L., and Ceccarelli E. A. (2014). Recombinant protein expression in Escherichia coli: advances and challenges. Front. Microbiol. 5:172. DOI:10.3389/fmicb.2014.00172 |
| [66] | Roope L. S., Smith R. D., Pouwels K. B., et al. (2019). The challenge of antimicrobial resistance: what economics can contribute. Science 364:eaau4679. DOI:10.1126/science.aau4679 |
| [67] | Bahar A. A. and Ren D. (2013). Antimicrobial peptides. Pharmaceuticals 6:1543−1575. DOI:10.3390/ph6121543 |
| [68] | Chen X., Li C., Bernards M. T., et al. (2021). Sequence-based peptide identification, generation, and property prediction with deep learning: A review. Mol. Syst. Des. Eng. 6:406−428. DOI:10.1039/D0ME00161A |
| [69] | Yan J., Cai J., Zhang B., et al. (2022). Recent progress in the discovery and design of antimicrobial peptides using traditional machine learning and deep learning. Antibiotics 11:1451. DOI:10.3390/antibiotics11101451 |
| [70] | Ramazi S., Mohammadi N., Allahverdi A., et al. (2022). A review on antimicrobial peptides databases and the computational tools. Database 2022:baac011. DOI:10.1093/database/baac011 |
| [71] | Capecchi A., Cai X., Personne H., et al. (2021). Machine learning designs non-hemolytic antimicrobial peptides. Chem. Sci. 12:9221−9232. DOI:10.1039/D1SC01713F |
| [72] | Dean S. N., Alvarez J. A. E., Zabetakis D., et al. (2021). PepVAE: variational autoencoder framework for antimicrobial peptide generation and activity prediction. Front. Microbiol. 12:725727. DOI:10.3389/fmicb.2021.725727 |
| [73] | Das P., Sercu T., Wadhawan K., et al. (2021). Accelerated antimicrobial discovery via deep generative models and molecular dynamics simulations. Nat. Biomed. Eng. 5:613−623. DOI:10.1038/s41551-021-00689-x |
| [74] | Tucs A., Tran D. P., Yumoto A., et al. (2020). Generating ampicillin-level antimicrobial peptides with activity-aware generative adversarial networks. ACS Omega 5:22847−22851. DOI:10.1021/acsomega.0c02088 |
| [75] | Jang W. D., Kim G. B., Kim Y., et al. (2022). Applications of artificial intelligence to enzyme and pathway design for metabolic engineering. Curr. Opin. Biotechnol. 73:101−107. DOI:10.1016/j.copbio.2021.07.024 |
| [76] | Pleiss J. (2011). Protein design in metabolic engineering and synthetic biology. Curr. Opin. Biotechnol. 22:611−617. DOI:10.1016/j.copbio.2011.03.004 |
| [77] | Khurana S., Rawi R., Kunji K., et al. (2018). DeepSol: a deep learning framework for sequence-based protein solubility prediction. Bioinformatics 34:2605−2613. DOI:10.1093/bioinformatics/bty166 |
| [78] | Han X., Wang X., and Zhou K. (2019). Develop machine learning-based regression predictive models for engineering protein solubility. Bioinformatics 35:4640−4646. DOI:10.1093/bioinformatics/btz294 |
| [79] | Li F., Yuan L., Lu H., et al. (2022). Deep learning-based k cat prediction enables improved enzyme-constrained model reconstruction. Nat. Catal. 5:662−672. DOI:10.1038/s41929-022-00798-z |
| [80] | Yu H., and Luo X. (2022). Pretrained language models and weight redistribution achieve precise kcat prediction. bioRxiv:2022.2011. 2023.517595. DOI:10.1101/2022.11.23.517595. |
| [81] | Hu X., Feng C., Ling T., et al. (2022). Deep learning frameworks for protein-protein interaction prediction. Comput. Struct. Biotechnol. J. 20:3223−3233. DOI:10.1016/j.csbj.2022.06.025 |
| [82] | Casadio R., Martelli P. L., and Savojardo C. (2022). Machine learning solutions for predicting protein–protein interactions. Wiley Interdiscip. Rev.:Comput. Mol. Sci. 12:e1618. DOI:10.1002/wcms.1618 |
| [83] | Yeh A. H.-W., Norn C., Kipnis Y., et al. (2023). De novo design of luciferases using deep learning. Nature 614:774−780. DOI:10.1038/s41586-023-05696-3 |
| [84] | Watson J. L., Juergens D., Bennett N. R., et al. (2023). De novo design of protein structure and function with RFdiffusion. Nature 620:1089−1100. DOI:10.1038/s41586-023-05696-3 |
| [85] | Yu L., Zhang W., Wang J., et al. (2017). Seqgan: Sequence generative adversarial nets with policy gradient. Proceedings of the AAAI conference on artificial intelligence.DOI:10.1609/aaai.v31i1.10804 |
| [86] | Arjovsky M., and Bottou L. (2017). Towards principled methods for training generative adversarial networks. arXiv preprint arXiv:1701.04862. DOI:10.48550/arXiv.1701.04862 |
| [87] | Caccia M., Caccia L., Fedus W., et al. (2018). Language gans falling short. arXiv preprint arXiv:1811.02549. DOI:10.48550/arXiv.1811.02549 |
| [88] | Thanh-Tung H., and Tran T. (2020). Catastrophic forgetting and mode collapse in gans. 2020 international joint conference on neural networks (ijcnn). IEEE. DOI:10.1109/IJCNN48605.2020.9207181 |
| [89] | Shmelkov K., Schmid C., and Alahari K. (2018). How good is my GAN? Proceedings of the European conference on computer vision (ECCV), pp:218 - 234. DOI:10.1007/978-3-030-01216-8_14 |
| [90] | Szymczak P., Możejko M., Grzegorzek T., et al. (2023). Discovering highly potent antimicrobial peptides with deep generative model HydrAMP. Nat. Commun. 14:1453. DOI:10.1038/s41467-023-36994-z |
| Zeng Z., Xu R., Guo J., et al. (2025). Accelerating functional protein discovery with GPT models: Antimicrobials and enzymes. The Innovation Life 3:100133. https://doi.org/10.59717/j.xinn-life.2025.100133 |
To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.
Discriminator (A) and generator (B) architecture
Accuracies (A) and resolutions (B) of the discriminators
Effects of generator finetuning on AMP (A) and MDH (B) likelihood of the generated sequences
Computational analysis and evidence for the functionality of the prioritized MDH candidates
Experimental validation on the generated and prioritized MDH candidates