The study explores the security of Artificial Intelligence peer-review systems in academic evaluation.
Experiments reveal AI peer-review systems favor prestigious institutions over unknown ones.
AI peer-review systems penalize cautious writing and yield to confident but evidence-free rebuttals.
AI peer-review systems are also vulnerable to poisoning attacks using manipulated data.
| [1] | Sample I. (2025). Quality of scientific papers questioned as academics “overwhelmed” by the millions published. The Guardian. https://www.theguardian.com/science/2025/jul/13/quality-of-scientific-papers-questioned-as-academics-overwhelmed-by-the-millions-published |
| [2] | Adam D. (2025). The peer-review crisis: How to fix an overloaded system. Nature 644:24−27. DOI:10.1038/d41586-025-02457-2 |
| [3] | Bergstrom C.T. and Bak-Coleman J. (2025). AI, peer review and the human activity of science. Nature. DOI:10.1038/d41586-025-01839-w |
| [4] | Khalifa M. and Albadawy M. (2024). Using artificial intelligence in academic writing and research: An essential productivity tool. Comput. Methods Programs Biomed. Update 5:100145. DOI:10.1016/j.cmpbup.2024.100145 |
| [5] | Chen Q., Yang M., Qin L., et al. (2025). AI4Research: A survey of artificial intelligence for scientific research. arXiv preprint. DOI:10.48550/arXiv.2507.01903 |
| [6] | Luo Z., Yang Z., Xu Z., et al. (2025). Llm4sr: A survey on large language models for scientific research. arXiv preprint. DOI:10.48550/arXiv:2501.04306 |
| [7] | Liang W., Izzo Z., Zhang Y., et al. (2024). Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews. Proc. Int. Conf. Mach. Learn. 235:1192. DOI:10.5555/3692070.3693262 |
| [8] | Wu D. (2025). Researchers are using AI for peer reviews—and finding ways to cheat it. The Washington Post. https://www.washingtonpost.com/nation/2025/07/17/ai-university-research-peer-review/ |
| [9] | Tong T., Wang F., Zhao Z., et al. (2025). Badjudge: Backdoor vulnerabilities of llm-as-a-judge. arXiv preprint. DOI:10.48550/arXiv.2503.00596 |
| [10] | Gibney E. (2025). Scientists hide messages in papers to game AI peer review. Nature 643:887−888. DOI:10.1038/d41586-025-02172-y |
| [11] | Ji Z., Lee N., Frieske R., et al. (2023). Survey of hallucination in natural language generation. ACM Comput. Surv. 55:1−38. DOI:10.1145/3571730 |
| [12] | Jin Y., Zhao Q., Wang Y., et al. (2024). Agentreview: Exploring peer review dynamics with llm agents. arXiv preprint. DOI:10.48550/arXiv.2406.12708 |
| [13] | Ye J., Wang Y., Huang Y., et al. (2024). Justice or prejudice? quantifying biases in llm-as-a-judge. arXiv preprint. DOI:10.48550/arXiv.2410.02736 |
| [14] | Lin T.-L., Chen W.-C., Hsiao T.-F., et al. (2025). Breaking the reviewer: Assessing the vulnerability of large language models in automated peer review under textual adversarial attacks. arXiv preprint. DOI:10.48550/arXiv.2506.11113 |
| [15] | Li Y., Jiang Y., Li Z., et al. (2024). Backdoor learning: A survey. IEEE Trans. Neural Netw. Learn. Syst. 35:5−22. DOI:10.1109/TNNLS.2022.3182979 |
| [16] | Zhang Y., Rando J., Evtimov I., et al. (2024). Persistent pre-training poisoning of llms. arXiv preprint. DOI:10.48550/arXiv.2410.13722 |
| [17] | Perez F. and Ribeiro I. (2022). Ignore previous prompt: Attack techniques for language models. arXiv preprint. DOI:10.48550/arXiv.2211.09527 |
| [18] | Shayegani E., Mamun M.A.A., Fu Y., et al. (2023). Survey of vulnerabilities in large language models revealed by adversarial attacks. arXiv preprint. DOI:10.48550/arXiv.2310.10844 |
| [19] | Sharma M., Tong M., Korbak T., et al. (2023). Towards understanding sycophancy in language models. arXiv preprint. DOI:10.48550/arXiv.2310.13548 |
| [20] | Fanous A., Goldberg J., Agarwal A., et al. (2025). Syceval: Evaluating llm sycophancy. Proc. AAAI/ACM Conf. AI Ethics Soc. 8:893−900. DOI:10.48550/arXiv.2502.08177 |
| [21] | Shi J., Yuan Z., Liu Y., et al. (2024). Optimization-based prompt injection attack to llm-as-a-judge. Proc. ACM SIGSAC Conf. Comput. Commun. Secur. 2024:660–674. DOI:10.1145/3658644.3690291 |
| [22] | Malmqvist L. (2025). Sycophancy in large language models: Causes and mitigations. Intell. Comput. Proc. Comput. Conf. 2024:61−74. DOI:10.1007/978-3-031-92611-2_5 |
| [23] | Xu R., Lin B., Yang S., et al. (2024). The earth is flat because.: Investigating llms’ belief towards misinformation via persuasive conversation. Proc. Annu. Meet. Assoc. Comput. Linguist. 2024:16259−16303. DOI:10.18653/v1/2024.acl-long.858 |
| [24] | Nuijten M.B., van Assen M.A.L.M., Hartgerink C.H.J., et al. (2017). The validity of the tool “statcheck” in discovering statistical reporting inconsistencies. PsyArXiv preprint. DOI:10.31234/osf.io/tcxaj |
| [25] | Checco A., Bracciale L., Loreti P., et al. (2021). AI-assisted peer review. Humanit. Soc. Sci. Commun. 8:1−11. DOI:10.1057/s41599-020-00703-8 |
| [26] | Charlin L. and Zemel R.S. (2013). The toronto paper matching system: An automated paper-reviewer assignment system. ICML PEER. https://www.cs.toronto.edu/~lcharlin/papers/tpms.pdf |
| [27] | Leyton-Brown K., Mausam., Nandwani Y., et al. (2024). Matching papers and reviewers at large conferences. Artif. Intell. 331:104119. DOI:10.1016/j.artint.2024.104119 |
| [28] | Liu R. and Shah N.B. (2023). Reviewergpt? An exploratory study on using large language models for paper reviewing. arXiv preprint. DOI:10.48550/arXiv.2306.00622 |
| [29] | Gao Z., Brantley K. and Joachims T. (2024). Reviewer2: Optimizing review generation through prompt generation. arXiv preprint. DOI:10.48550/arXiv.2402.10886 |
| [30] | Yu J., Ding Z., Tan J., et al. (2024). Automated peer reviewing in paper sea: Standardization, evaluation, and analysis. arXiv preprint. DOI:10.18653/v1/2024.findings-emnlp.595 |
| [31] | Wang Q., Zeng Q., Huang L., et al. (2020). ReviewRobot: Explainable paper review generation based on knowledge synthesis. arXiv preprint. DOI:10.18653/v1/2020.inlg-1.44 |
| [32] | Weng Y., Zhu M., Bao G., et al. (2024). Cycleresearcher: Improving automated research via automated review. arXiv preprint. DOI:10.48550/arXiv.2411.00816 |
| [33] | D'Arcy M., Hope T., Birnbaum L., et al. (2024). Marg: Multi-agent review generation for scientific papers. arXiv preprint. DOI:10.48550/arXiv.2401.04259 |
| [34] | Taechoyotin P., Wang G., Zeng T., et al. (2024). MAMORX: Multi-agent multi-modal scientific review generation with external knowledge. Proc. NeurIPS Workshop Found. Models Sci. https://openreview.net/forum?id=frvkE8rCfX |
| [35] | Sun L., Chan A., Chang Y.S., et al. (2024). ReviewFlow: Intelligent scaffolding to support academic peer reviewing. Proc. Int. Conf. Intell. User Interfaces 2024:120−137. DOI:10.1145/3640543.3645159 |
| [36] | Zyska D., Dycke N., Buchmann J., et al. (2023). CARE: Collaborative AI-assisted reading environment. arXiv preprint. DOI:10.18653/v1/2023.acl-demo.28 |
| [37] | Mathur P., Siu A., Manjunatha V., et al. (2024). DocPilot: Copilot for automating PDF edit workflows in documents. Proc. Annu. Meet. Assoc. Comput. Linguist. 3:232−246. DOI:10.18653/v1/2024.acl-demos.22 |
| [38] | Shanahan D. (2016). A peerless review? Automating methodological and statistical review.https://blogs.biomedcentral.com/bmcblog/2016/05/23/peerless-review-automating-methodological-statistical-review/ |
| [39] | Cyranoski D. (2019). Artificial intelligence is selecting grant reviewers in China. Nature 569:316−317. DOI:10.1038/d41586-019-01517-8 |
| [40] | Skarlinski M.D., Cox S., Laurent J.M., et al. (2024). Language agents achieve superhuman synthesis of scientific knowledge. arXiv preprint. DOI:10.48550/arXiv.2409.13740 |
| [41] | Lin E., Peng Z. and Fang Y. (2025). Evaluating and enhancing large language models for novelty assessment in scholarly publications. Proc. Workshop AI Sci. Discov. 2025:46−57. DOI:10.18653/v1/2025.aisd-main.5 |
| [42] | Radensky M., Shahid S., Fok R., et al. (2024). Scideator: Human-llm scientific idea generation grounded in research-paper facet recombination. arXiv preprint. DOI:10.48550/arXiv.2409.14634 |
| [43] | Couto P.H., Ho Q.P., Kumari N., et al. (2024). Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance. arXiv preprint. DOI:10.48550/arXiv.2406.10294 |
| [44] | Faizullah A.R.B.M., Urlana A. and Mishra R. (2024). Limgen: Probing the llms for generating suggestive limitations of research papers. Proc. Jt. Eur. Conf. Mach. Learn. Knowl. Discov. Databases 2024:106−124. DOI:10.1007/978-3-031-70344-7_7 |
| [45] | Bhatia C., Pradhan T. and Pal S. (2020). Metagen: An academic meta-review generation system. Proc. Int. ACM SIGIR Conf. Res. Dev. Inf. Retr. 2020:1653−1656. DOI:10.1145/3397271.3401190 |
| [46] | Shen C., Cheng L., Zhou R., et al. (2022). MReD: A meta-review dataset for structure-controllable text generation. Findings Assoc. Comput. Linguist. ACL 2022:2521−2535. DOI:10.18653/v1/2022.findings-acl.198 |
| [47] | Zeng Q., Sidhu M., Blume A., et al. (2024). Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation. Proc. Int. Jt. Conf. Artif. Intell. 2024:20−38. DOI:10.1007/978-981-97-9536-9_2 |
| [48] | Li M., Hovy E. and Lau J. (2023). Summarizing multiple documents with conversational structure for meta-review generation. Findings Assoc. Comput. Linguist. EMNLP 2023:7089−7112. DOI:10.18653/v1/2023.findings-emnlp.472 |
| [49] | Sun L., Tao S., Hu J., et al. (2024). MetaWriter: Exploring the the potential and perils of ai writing support in scientific peer review. Proc. ACM Hum.-Comput. Interact. 8:1−32. DOI:10.1145/3637371 |
| [50] | Darrin M., Arous I., Piantanida P., et al. (2024). Glimpse: Pragmatically informative multi-document summarization for scholarly reviews. Proc. Annu. Meet. Assoc. Comput. Linguist. 2024:12737−12752. DOI:10.18653/v1/2024.acl-long.688 |
| [51] | Sukpanichnant P., Rapberger A. and Toni F. (2024). Peerarg: Argumentative peer review with llms. arXiv preprint. DOI:10.48550/arXiv.2409.16813 |
| [52] | Hossain E., Sinha S.K., Bansal N., et al. (2025). Llms as meta-reviewers’ assistants: A case study. Proc. Conf. North Am. Chapter Assoc. Comput. Linguist. Hum. Lang. Technol. 2025:7763−7803. DOI:10.18653/v1/2025.naacl-long.395 |
| [53] | Krizhevsky A., Sutskever I. and Hinton G.E. (2017). ImageNet classification with deep convolutional neural networks. Commun. ACM 60:84−90. DOI:10.1145/3065386 |
| [54] | Hinton G., Deng L., Yu D., et al. (2012). Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Process. Mag. 29:82−97. DOI:10.1109/msp.2012.2205597 |
| [55] | Devlin J., Chang M.-W., Lee K., et al. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. Proc. NAACL HLT:4171–4186. DOI:10.18653/v1/N19-1423 |
| [56] | Szegedy C., Zaremba W., Sutskever I., et al. (2013). Intriguing properties of neural networks. arXiv preprint. DOI:10.48550/arXiv.1312.6199 |
| [57] | Biggio B. and Roli F. (2018). Wild patterns: Ten years after the rise of adversarial machine learning. Proc. ACM SIGSAC Conf. Comput. Commun. Secur. 2018:2154−2156. DOI:10.1145/3243734.3264418 |
| [58] | Goodfellow I.J., Shlens J. and Szegedy C. (2014). Explaining and harnessing adversarial examples. arXiv preprint. DOI:10.48550/arXiv.1412.6572 |
| [59] | Athalye A., Carlini N. and Wagner D. (2018). Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. Proc. Int. Conf. Mach. Learn. 2018:274−283. DOI:10.48550/arXiv.1802.00420 |
| [60] | Barreno M., Nelson B., Sears R., et al. (2006). Can machine learning be secure? Proc. ACM Symp. Inf. Comput. Commun. Secur. 2006:16−25. DOI:10.1145/1128817.1128824 |
| [61] | Biggio B., Corona I., Maiorca D., et al. (2013). Evasion attacks against machine learning at test time. Proc. Jt. Eur. Conf. Mach. Learn. Knowl. Discov. Databases 2013:387−402. DOI:10.1007/978-3-642-40994-3_25 |
| [62] | Carlini N. and Wagner D. (2017). Towards evaluating the robustness of neural networks. Proc. IEEE Symp. Secur. Priv. 2017:39−57. DOI:10.1109/SP.2017.49 |
| [63] | Papernot N., McDaniel P., Jha S., et al. (2016). The limitations of deep learning in adversarial settings. Proc. IEEE Eur. Symp. Secur. Priv. 2016:372−387. DOI:10.1109/EuroSP.2016.36 |
| [64] | Madry A., Makelov A., Schmidt L., et al. (2017). Towards deep learning models resistant to adversarial attacks. arXiv preprint. DOI:10.48550/arXiv.1706.06083 |
| [65] | Chen P.-Y., Zhang H., Sharma Y., et al. (2017). Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. Proc. ACM Workshop Artif. Intell. Secur. 2017:15−26. DOI:10.1145/3128572.3140448 |
| [66] | Ilyas A., Engstrom L., Athalye A., et al. (2018). Black-box adversarial attacks with limited queries and information. Proc. Int. Conf. Mach. Learn. 2018:2137−2146. DOI:10.48550/arXiv.1804.08598 |
| [67] | Papernot N., McDaniel P., Sinha A., et al. (2018). Sok: Security and privacy in machine learning. Proc. IEEE Eur. Symp. Secur. Priv. 2018:399−414. DOI:10.1109/EuroSP.2018.00035 |
| [68] | Fredrikson M., Jha S. and Ristenpart T. (2015). Model inversion attacks that exploit confidence information and basic countermeasures. Proc. ACM SIGSAC Conf. Comput. Commun. Secur. 2015:1322−1333. DOI:10.1145/2810103.2813677 |
| [69] | Shokri R., Stronati M., Song C., et al. (2017). Membership inference attacks against machine learning models. Proc. IEEE Symp. Secur. Priv. 2017:3−18. DOI:10.1109/SP.2017.41 |
| [70] | Tramèr F., Zhang F., Juels A., et al. (2016). Stealing machine learning models via prediction APIs. Proc. USENIX Secur. Symp. 2016:601−618. DOI:10.5555/3241094.3241142 |
| [71] | Yeom S., Giacomelli I., Fredrikson M., et al. (2018). Privacy risk in machine learning: Analyzing the connection to overfitting. Proc. IEEE Comput. Secur. Found. Symp. 2018:268−282. DOI:10.1109/CSF.2018.00027 |
| [72] | Biggio B., Nelson B. and Laskov P. (2012). Poisoning attacks against support vector machines. arXiv preprint. DOI:10.5555/3042573.3042761 |
| [73] | Tolpegin V., Truex S., Gursoy M.E., et al. (2020). Data poisoning attacks against federated learning systems. Proc. Eur. Symp. Res. Comput. Secur. 2020:480−501. DOI:10.1007/978-3-030-58951-6_24 |
| [74] | Gu T., Dolan-Gavitt B. and Garg S. (2017). Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint. DOI:10.48550/arXiv.1708.06733 |
| [75] | Chen X., Liu C., Li B., et al. (2017). Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint. DOI:10.48550/arXiv.1712.05526 |
| [76] | Shafahi A., Huang W.R., Najibi M., et al. (2018). Poison frogs! Targeted clean-label poisoning attacks on neural networks. Adv. Neural Inf. Process. Syst. 31:6106−6116. DOI:10.5555/3327345.3327509 |
| [77] | Zhang J., Chen B., Cheng X., et al. (2021). PoisonGAN: Generative poisoning attacks against federated learning in edge computing systems. IEEE Internet Things J. 8:3310−3322. DOI:10.1109/jiot.2020.3023126 |
| [78] | Carlini N., Athalye A., Papernot N., et al. (2019). On evaluating adversarial robustness. arXiv preprint. DOI:10.48550/arXiv.1902.06705 |
| [79] | Tramèr F., Kurakin A., Papernot N., et al. (2017). Ensemble adversarial training: Attacks and defenses. arXiv preprint. DOI:10.48550/arXiv.1705.07204 |
| [80] | Cohen J., Rosenfeld E. and Kolter Z. (2019). Certified adversarial robustness via randomized smoothing. Proc. Int. Conf. Mach. Learn. 2019:1310−1320. DOI:10.48550/arXiv.1902.02918 |
| [81] | Wu D., Xia S.-T. and Wang Y. (2020). Adversarial weight perturbation helps robust generalization. Adv. Neural Inf. Process. Syst. 33:2958−2969. DOI:10.5555/3495724.3495973 |
| [82] | Chen T., Liu S., Chang S., et al. (2020). Adversarial robustness: From self-supervised pre-training to fine-tuning. Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. 2020:699−708. DOI:10.1109/CVPR42600.2020.00078 |
| [83] | Metzen J.H., Genewein T., Fischer V., et al. (2017). On detecting adversarial perturbations. arXiv preprint. DOI:10.48550/arXiv.1702.04267 |
| [84] | Steinhardt J., Koh P.W.W. and Liang P.S. (2017). Certified defenses for data poisoning attacks. Adv. Neural Inf. Process. Syst. 30:3520−3532. DOI:10.5555/3294996.3295110 |
| [85] | Piet J., Alrashed M., Sitawarin C., et al. (2024). Jatmo: Prompt injection defense by task-specific finetuning. Proc. Eur. Symp. Res. Comput. Secur. 2024:105−124. DOI:10.1007/978-3-031-70879-4_6 |
| [86] | Doskaliuk B., Zimba O., Yessirkepov M., et al. (2025). Artificial Intelligence in peer review: Enhancing efficiency while preserving integrity. J. Korean Med. Sci. 40:e92. DOI:10.3346/jkms.2025.40.e92https://www.ncbi.nlm.nih.gov/pubmed/39995259 |
| [87] | Mann S.P., Aboy M., Seah J.J., et al. (2025). AI and the future of academic peer review. arXiv preprint. DOI:10.48550/arXiv.2509.14189 |
| [88] | Maturo F., Porreca A. and Porreca A. (2025). The risks of artificial intelligence in research: Ethical and methodological challenges in the peer review process. AI Ethics 5:5389−5396. DOI:10.1007/s43681-025-00775-9 |
| [89] | Keuper J. (2025). Prompt injection attacks on llm generated reviews of scientific publications. arXiv preprint. DOI:10.48550/arXiv.2509.10248 |
| [90] | Zhao Y., Liu H., Yu D., et al. (2025). One token to fool llm-as-a-judge. arXiv preprint. DOI:10.48550/arXiv.2507.08794 |
| [91] | Collu M.G., Salviati U., Confalonieri R., et al. (2025). Publish to perish: Prompt injection attacks on llm-assisted peer review. arXiv preprint. DOI:10.48550/arXiv.2508.20863 |
| [92] | Ye R., Pang X., Chai J., et al. (2024). Are we there yet? Revealing the risks of utilizing large language models in scholarly peer review. arXiv preprint. DOI:10.48550/arXiv.2412.01708 |
| [93] | Dong Y., Jiang X., Liu H., et al. (2024). Generalization or memorization: Data contamination and trustworthy evaluation for large language models. arXiv preprint. DOI:10.18653/v1/2024.findings-acl.716 |
| [94] | Goldblum M., Tsipras D., Xie C., et al. (2023). Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses. IEEE Trans. Pattern Anal. Mach. Intell. 45:1563−1580. DOI:10.1109/TPAMI.2022.3162397 |
| [95] | Borgeaud S., Mensch A., Hoffmann J., et al. (2022). Improving language models by retrieving from trillions of tokens. Proc. Int. Conf. Mach. Learn. 2022:2206−2240. DOI:10.48550/arXiv.2112.04426 |
| [96] | Souly A., Rando J., Chapman E., et al. (2025). Poisoning attacks on llms require a near-constant number of poison samples. arXiv preprint. DOI:10.48550/arXiv.2510.07192 |
| [97] | Wen J., Si C., Chen Y.-h., et al. (2025). Predicting empirical ai research outcomes with language models. arXiv preprint. DOI:10.48550/arXiv.2506.00794 |
| [98] | Bereska L. and Gavves E. (2024). Mechanistic interpretability for AI safety-A review. arXiv preprint. DOI:10.48550/arXiv.2404.14082 |
| [99] | Lo L.Y. and Qu H. (2025). How good (or bad) are llms at detecting misleading visualizations? IEEE Trans. Vis. Comput. Graph. 31:1116−1125. DOI:10.1109/TVCG.2024.3456333 |
| [100] | Tonglet J., Zimny J., Tuytelaars T., et al. (2025). Is this chart lying to me? Automating the detection of misleading visualizations. arXiv preprint. DOI:10.48550/arXiv.2508.21675 |
| [101] | Gallegos I.O., Rossi R.A., Barrow J., et al. (2024). Bias and fairness in large language models: A survey. Comput. Linguist. 50:1097−1179. DOI:10.1162/coli_a_00524 |
| [102] | Liu Y., Deng G., Li Y., et al. (2023). Prompt injection attack against llm-integrated applications. arXiv preprint. DOI:10.48550/arXiv.2306.05499 |
| [103] | Zhou X., Qiang Y., Zade S.Z., et al. (2023). Hijacking large language models via adversarial in-context learning. arXiv preprint. DOI:10.48550/arXiv.2311.09948 |
| [104] | Gong Y., Chen Z., Chen M., et al. (2025). Topic-fliprag: Topic-orientated adversarial opinion manipulation attacks to retrieval-augmented generation models. arXiv preprint. DOI:10.5555/3766078.3766274 |
| [105] | Schwinn L., Dobre D., Günnemann S., et al. (2023). Adversarial attacks and defenses in large language models: Old and new threats. arXiv preprint. DOI:10.48550/arxiv.2310.19737 |
| [106] | Raina V., Liusie A. and Gales M. (2024). Is llm-as-a-judge robust? Investigating universal adversarial attacks on zero-shot llm assessment. arXiv preprint. DOI:10.48550/arXiv.2402.14016 |
| [107] | Guo Y., Guo M., Su J., et al. (2024). Bias in large language models: Origin, evaluation, and mitigation. arXiv preprint. DOI:10.48550/arXiv.2411.10915 |
| [108] | Navigli R., Conia S. and Ross B. (2023). Biases in large language models: Origins, inventory, and discussion. ACM J. Data Inf. Qual. 15:1−21. DOI:10.1145/3597307 |
| [109] | Angrist J.D. (2014). The perils of peer effects. Lab. Econ. 30:98−108. DOI:10.1016/j.labeco.2014.05.008 |
| [110] | Lewis P., Perez E., Piktus A., et al. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Adv. Neural Inf. Process. Syst. 33:9459−9474. DOI:10.5555/3495724.3496517 |
| [111] | Schwarzschild A., Goldblum M., Gupta A., et al. (2021). Just how toxic is data poisoning. a unified benchmark for backdoor and data poisoning attacks. Proc. Int. Conf. Mach. Learn. 2021:9389−9398. DOI:10.48550/arXiv.2006.12557 |
| [112] | Touvron H., Lavril T., Izacard G., et al. (2023). Llama: Open and efficient foundation language models. arXiv preprint. DOI:10.48550/arXiv.2302.13971 |
| [113] | Bowen D., Murphy B., Cai W., et al. (2025). Scaling trends for data poisoning in LLMs. Proc. AAAI Conf. Artif. Intell. 39:27206−27214. DOI:10.1609/aaai.v39i26.34929 |
| [114] | Liu T., Zhang Y., Feng Z., et al. (2024). Beyond traditional threats: A persistent backdoor attack on federated learning. Proc. AAAI Conf. Artif. Intell. 38:21359−21367. DOI:10.1609/aaai.v38i19.30131 |
| [115] | Zhu C., Li Y., Rao B., et al. (2025). SPA: Towards more stealth and persistent backdoor attacks in federated learning. arXiv preprint. DOI:10.48550/arXiv.2506.20931 |
| [116] | Tian Z., Cui L., Liang J., et al. (2022). A comprehensive survey on poisoning attacks and countermeasures in machine learning. ACM Comput. Surv. 55:1−35. DOI:10.1145/3551636 |
| [117] | Zhao P., Zhu W., Jiao P., et al. (2025). Data poisoning in deep learning: A survey. arXiv preprint. DOI:10.48550/arXiv.2503.22759 |
| [118] | Muñoz-González L., Biggio B., Demontis A., et al. (2017). Towards poisoning of deep learning algorithms with back-gradient optimization. Proc. ACM Workshop Artif. Intell. Secur. 2017:27−38. DOI:10.1145/3128572.3140451 |
| [119] | Nourani M., Roy C., Block J.E., et al. (2021). Anchoring bias affects mental model formation and user reliance in explainable AI systems. Proc. Int. Conf. Intell. User Interfaces 2021:340−350. DOI:10.1145/3397481.3450639 |
| [120] | Shi F., Chen X., Misra K., et al. (2023). Large language models can be easily distracted by irrelevant context. Proc. Int. Conf. Mach. Learn. 2023:31210−31227. DOI:10.5555/3618408.3619699 |
| [121] | Dougrez-Lewis J., Akhter M.E., Ruggeri F., et al. (2025). Assessing the the reasoning capabilities of llms in the context of evidence-based claim verification. Findings Assoc. Comput. Linguist. ACL 2025:20604−20628. DOI:10.18653/v1/2025.findings-acl.1059 |
| [122] | Hong R., Zhang H., Pang X., et al. (2024). A closer look at the self-verification abilities of large language models in logical reasoning. Proc. NAACL HLT 2024:900−925. DOI:10.18653/v1/2024.naacl-long.52 |
| [123] | Li J., Li Y., Hu X., et al. (2025). Aspect-guided multi-level perturbation analysis of large language models in automated peer review. arXiv preprint. DOI:10.48550/arXiv.2502.12510 |
| [124] | Liang W., Zhang Y., Cao H., et al. (2024). Can large language models provide useful feedback on research papers? A large-scale empirical analysis. NEJM AI 1:AIoa2400196. DOI:10.1056/AIoa2400196 |
| [125] | Zhou Z., Li Z., Zhang J., et al. (2025). Corba: Contagious recursive blocking attacks on multi-agent systems based on large language models. arXiv preprint. DOI:10.48550/arXiv.2502.14529 |
| [126] | Zhu S., Zhang R., An B., et al. (2023). Autodan: Interpretable gradient-based adversarial attacks on large language models. arXiv preprint. DOI:10.48550/arXiv.2310.15140 |
| [127] | Zizzo G., Cornacchia G., Fraser K., et al. (2025). Adversarial prompt evaluation: Systematic benchmarking of guardrails against prompt input attacks on llms. arXiv preprint. DOI:10.48550/arXiv.2502.15427 |
| [128] | Bozdag N.B., Mehri S., Tur G., et al. (2025). Persuade me if you can: A framework for evaluating persuasion effectiveness and susceptibility among large language models. arXiv preprint. DOI:10.48550/arXiv.2503.01829 |
| [129] | Salvi F., Horta Ribeiro M., Gallotti R., et al. (2025). On the conversational persuasiveness of GPT-4. Nat. Hum. Behav. 9:1645−1653. DOI:10.1038/s41562-025-02194-6 |
| [130] | Liu Y., Yang K., Liu Y., et al. (2023). The shackles of peer review: Unveiling the flaws in the ivory tower. arXiv preprint. DOI:10.48550/arXiv.2310.05966 |
| [131] | Nisbett R.E. and Wilson T.D. (1977). The halo effect: Evidence for unconscious alteration of judgments. J. Pers. Soc. Psychol. 35:250−256. DOI:10.1037/0022-3514.35.4.250 |
| [132] | Zhang J., Zhang H., Deng Z., et al. (2022). Investigating fairness disparities in peer review: A language model enhanced approach. arXiv preprint. DOI:10.48550/arXiv.2211.06398 |
| [133] | Fox C.W., Meyer J. and Aimé E. (2023). Double‐blind peer review affects reviewer ratings and editor decisions at an ecology journal. Funct. Ecol. 37:1144−1157. DOI:10.1111/1365-2435.14259 |
| [134] | Sun M., Barry Danfa J. and Teplitskiy M. (2021). Does double‐blind peer review reduce bias. Evidence from a top computer science conference. J. Assoc. Inf. Sci. Technol. 73:811−819. DOI:10.1002/asi.24582 |
| [135] | Verharen J.P.H. (2023). ChatGPT identifies gender disparities in scientific peer review. eLife 12:RP90230. DOI:10.7554/eLife.90230 |
| [136] | Hosseini M. and Horbach S. (2023). Fighting reviewer fatigue or amplifying bias. Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer review. Res. Integr. Peer Rev. 8:4. DOI:10.1186/s41073-023-00133-5 |
| [137] | Soneji A., Kokulu F.B., Rubio-Medrano C., et al. (2022). “Flawed, but like democracy we don’t have a better system”: The Experts’ Insights on the peer review process of evaluating security papers. Proc. IEEE Symp. Secur. Priv. 2022:1845−1862. DOI:10.1109/SP46214.2022.9833581 |
| [138] | Schramowski P., Turan C., Andersen N., et al. (2022). Large pre-trained language models contain human-like biases of what is right and wrong to do. Nat. Mach. Intell. 4:258−268. DOI:10.1038/s42256-022-00458-8 |
| [139] | Li H., Shan S., Wenger E., et al. (2022). Blacklight: Scalable defense for neural networks against query-based black-box attacks. Proc. USENIX Secur. Symp. 2022:2117–2134. https://www.usenix.org/conference/usenixsecurity22/presentation/li-huiying |
| [140] | Koo R., Lee M., Raheja V., et al. (2024). Benchmarking cognitive biases in large language models as evaluators. Findings Assoc. Comput. Linguist. ACL 2024:517−545. DOI:10.18653/v1/2024.findings-acl.29 |
| [141] | Bartos O.J. and Wehr P. (2002). Using conflict theory (Cambridge University Press). DOI:10.2307/1556608 |
| [142] | Robertson Z. (2023). GPT4 is slightly helpful for peer-review assistance: A pilot study. arXiv preprint. DOI:10.48550/arXiv.2307.05492 |
| Wang J., Liu Y., Xu H., et al. (2026). When AI reviews science: Can we trust the referee? The Innovation Informatics 2:100030. https://doi.org/10.59717/j.xinn-inform.2026.100030 |
To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.
AI peer-review loop
A system vulnerability in the OpenReview platform led to the leakage of the identity information of reviewers and authors.
Overview of the threat model for an AI peer-review pipeline, detailing various attack methods and the specific stages they target.
Identity Bias Exploitation
Sensitivity to Assertion Strength
Sycophancy in the Rebuttal
Contextual Poisoning