Article Contents
REVIEW   Open Access     Cite

When AI reviews science: Can we trust the referee?

More Information
  • DownLoad: Full size image
    1. The study explores the security of Artificial Intelligence peer-review systems in academic evaluation.

      Experiments reveal AI peer-review systems favor prestigious institutions over unknown ones.

      AI peer-review systems penalize cautious writing and yield to confident but evidence-free rebuttals.

      AI peer-review systems are also vulnerable to poisoning attacks using manipulated data.

  • The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large language models (LLMs) offer impressive capabilities in summarization, fact checking, and literature triage, making the integration of AI into peer review increasingly attractive—and, in practice, unavoidable. Yet early deployments and informal adoption have exposed acute failure modes. Recent incidents have revealed that hidden prompt injections embedded in manuscripts can steer LLM-generated reviews toward unjustifiably positive judgments. Complementary studies have also demonstrated brittleness to adversarial phrasing, authority and length biases, and hallucinated claims. These episodes raise a central question for scholarly communication: when AI reviews science, can we trust the AI referee? This paper provides a security- and reliability-centered analysis of AI peer review. We map attacks across the review lifecycle—training and data retrieval, desk review, deep review, rebuttal, and system-level. We instantiate this taxonomy with four treatment-control probes on a stratified set of ICLR 2025 submissions, using two advanced LLM-based referees to isolate the causal effects of prestige framing, assertion strength, rebuttal sycophancy, and contextual poisoning on review scores. Together, this taxonomy and experimental audit provide an evidence-based baseline for assessing and tracking the reliability of AI peer review and highlight concrete failure points to guide targeted, testable mitigations.
  • 加载中
  • [1] Sample I. (2025). Quality of scientific papers questioned as academics “overwhelmed” by the millions published. The Guardian. https://www.theguardian.com/science/2025/jul/13/quality-of-scientific-papers-questioned-as-academics-overwhelmed-by-the-millions-published

    View in Article Google Scholar

    [2] Adam D. (2025). The peer-review crisis: How to fix an overloaded system. Nature 644:24−27. DOI:10.1038/d41586-025-02457-2

    View in Article CrossRef Google Scholar

    [3] Bergstrom C.T. and Bak-Coleman J. (2025). AI, peer review and the human activity of science. Nature. DOI:10.1038/d41586-025-01839-w

    View in Article Google Scholar

    [4] Khalifa M. and Albadawy M. (2024). Using artificial intelligence in academic writing and research: An essential productivity tool. Comput. Methods Programs Biomed. Update 5:100145. DOI:10.1016/j.cmpbup.2024.100145

    View in Article CrossRef Google Scholar

    [5] Chen Q., Yang M., Qin L., et al. (2025). AI4Research: A survey of artificial intelligence for scientific research. arXiv preprint. DOI:10.48550/arXiv.2507.01903

    View in Article Google Scholar

    [6] Luo Z., Yang Z., Xu Z., et al. (2025). Llm4sr: A survey on large language models for scientific research. arXiv preprint. DOI:10.48550/arXiv:2501.04306

    View in Article Google Scholar

    [7] Liang W., Izzo Z., Zhang Y., et al. (2024). Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews. Proc. Int. Conf. Mach. Learn. 235:1192. DOI:10.5555/3692070.3693262

    View in Article Google Scholar

    [8] Wu D. (2025). Researchers are using AI for peer reviews—and finding ways to cheat it. The Washington Post. https://www.washingtonpost.com/nation/2025/07/17/ai-university-research-peer-review/

    View in Article Google Scholar

    [9] Tong T., Wang F., Zhao Z., et al. (2025). Badjudge: Backdoor vulnerabilities of llm-as-a-judge. arXiv preprint. DOI:10.48550/arXiv.2503.00596

    View in Article Google Scholar

    [10] Gibney E. (2025). Scientists hide messages in papers to game AI peer review. Nature 643:887−888. DOI:10.1038/d41586-025-02172-y

    View in Article CrossRef Google Scholar

    [11] Ji Z., Lee N., Frieske R., et al. (2023). Survey of hallucination in natural language generation. ACM Comput. Surv. 55:1−38. DOI:10.1145/3571730

    View in Article CrossRef Google Scholar

    [12] Jin Y., Zhao Q., Wang Y., et al. (2024). Agentreview: Exploring peer review dynamics with llm agents. arXiv preprint. DOI:10.48550/arXiv.2406.12708

    View in Article Google Scholar

    [13] Ye J., Wang Y., Huang Y., et al. (2024). Justice or prejudice? quantifying biases in llm-as-a-judge. arXiv preprint. DOI:10.48550/arXiv.2410.02736

    View in Article Google Scholar

    [14] Lin T.-L., Chen W.-C., Hsiao T.-F., et al. (2025). Breaking the reviewer: Assessing the vulnerability of large language models in automated peer review under textual adversarial attacks. arXiv preprint. DOI:10.48550/arXiv.2506.11113

    View in Article Google Scholar

    [15] Li Y., Jiang Y., Li Z., et al. (2024). Backdoor learning: A survey. IEEE Trans. Neural Netw. Learn. Syst. 35:5−22. DOI:10.1109/TNNLS.2022.3182979

    View in Article CrossRef Google Scholar

    [16] Zhang Y., Rando J., Evtimov I., et al. (2024). Persistent pre-training poisoning of llms. arXiv preprint. DOI:10.48550/arXiv.2410.13722

    View in Article Google Scholar

    [17] Perez F. and Ribeiro I. (2022). Ignore previous prompt: Attack techniques for language models. arXiv preprint. DOI:10.48550/arXiv.2211.09527

    View in Article Google Scholar

    [18] Shayegani E., Mamun M.A.A., Fu Y., et al. (2023). Survey of vulnerabilities in large language models revealed by adversarial attacks. arXiv preprint. DOI:10.48550/arXiv.2310.10844

    View in Article Google Scholar

    [19] Sharma M., Tong M., Korbak T., et al. (2023). Towards understanding sycophancy in language models. arXiv preprint. DOI:10.48550/arXiv.2310.13548

    View in Article Google Scholar

    [20] Fanous A., Goldberg J., Agarwal A., et al. (2025). Syceval: Evaluating llm sycophancy. Proc. AAAI/ACM Conf. AI Ethics Soc. 8:893−900. DOI:10.48550/arXiv.2502.08177

    View in Article CrossRef Google Scholar

    [21] Shi J., Yuan Z., Liu Y., et al. (2024). Optimization-based prompt injection attack to llm-as-a-judge. Proc. ACM SIGSAC Conf. Comput. Commun. Secur. 2024:660–674. DOI:10.1145/3658644.3690291

    View in Article Google Scholar

    [22] Malmqvist L. (2025). Sycophancy in large language models: Causes and mitigations. Intell. Comput. Proc. Comput. Conf. 2024:61−74. DOI:10.1007/978-3-031-92611-2_5

    View in Article CrossRef Google Scholar

    [23] Xu R., Lin B., Yang S., et al. (2024). The earth is flat because.: Investigating llms’ belief towards misinformation via persuasive conversation. Proc. Annu. Meet. Assoc. Comput. Linguist. 2024:16259−16303. DOI:10.18653/v1/2024.acl-long.858

    View in Article CrossRef Google Scholar

    [24] Nuijten M.B., van Assen M.A.L.M., Hartgerink C.H.J., et al. (2017). The validity of the tool “statcheck” in discovering statistical reporting inconsistencies. PsyArXiv preprint. DOI:10.31234/osf.io/tcxaj

    View in Article Google Scholar

    [25] Checco A., Bracciale L., Loreti P., et al. (2021). AI-assisted peer review. Humanit. Soc. Sci. Commun. 8:1−11. DOI:10.1057/s41599-020-00703-8

    View in Article CrossRef Google Scholar

    [26] Charlin L. and Zemel R.S. (2013). The toronto paper matching system: An automated paper-reviewer assignment system. ICML PEER. https://www.cs.toronto.edu/~lcharlin/papers/tpms.pdf

    View in Article Google Scholar

    [27] Leyton-Brown K., Mausam., Nandwani Y., et al. (2024). Matching papers and reviewers at large conferences. Artif. Intell. 331:104119. DOI:10.1016/j.artint.2024.104119

    View in Article CrossRef Google Scholar

    [28] Liu R. and Shah N.B. (2023). Reviewergpt? An exploratory study on using large language models for paper reviewing. arXiv preprint. DOI:10.48550/arXiv.2306.00622

    View in Article Google Scholar

    [29] Gao Z., Brantley K. and Joachims T. (2024). Reviewer2: Optimizing review generation through prompt generation. arXiv preprint. DOI:10.48550/arXiv.2402.10886

    View in Article Google Scholar

    [30] Yu J., Ding Z., Tan J., et al. (2024). Automated peer reviewing in paper sea: Standardization, evaluation, and analysis. arXiv preprint. DOI:10.18653/v1/2024.findings-emnlp.595

    View in Article Google Scholar

    [31] Wang Q., Zeng Q., Huang L., et al. (2020). ReviewRobot: Explainable paper review generation based on knowledge synthesis. arXiv preprint. DOI:10.18653/v1/2020.inlg-1.44

    View in Article Google Scholar

    [32] Weng Y., Zhu M., Bao G., et al. (2024). Cycleresearcher: Improving automated research via automated review. arXiv preprint. DOI:10.48550/arXiv.2411.00816

    View in Article Google Scholar

    [33] D'Arcy M., Hope T., Birnbaum L., et al. (2024). Marg: Multi-agent review generation for scientific papers. arXiv preprint. DOI:10.48550/arXiv.2401.04259

    View in Article Google Scholar

    [34] Taechoyotin P., Wang G., Zeng T., et al. (2024). MAMORX: Multi-agent multi-modal scientific review generation with external knowledge. Proc. NeurIPS Workshop Found. Models Sci. https://openreview.net/forum?id=frvkE8rCfX

    View in Article Google Scholar

    [35] Sun L., Chan A., Chang Y.S., et al. (2024). ReviewFlow: Intelligent scaffolding to support academic peer reviewing. Proc. Int. Conf. Intell. User Interfaces 2024:120−137. DOI:10.1145/3640543.3645159

    View in Article CrossRef Google Scholar

    [36] Zyska D., Dycke N., Buchmann J., et al. (2023). CARE: Collaborative AI-assisted reading environment. arXiv preprint. DOI:10.18653/v1/2023.acl-demo.28

    View in Article Google Scholar

    [37] Mathur P., Siu A., Manjunatha V., et al. (2024). DocPilot: Copilot for automating PDF edit workflows in documents. Proc. Annu. Meet. Assoc. Comput. Linguist. 3:232−246. DOI:10.18653/v1/2024.acl-demos.22

    View in Article CrossRef Google Scholar

    [38] Shanahan D. (2016). A peerless review? Automating methodological and statistical review.https://blogs.biomedcentral.com/bmcblog/2016/05/23/peerless-review-automating-methodological-statistical-review/

    View in Article Google Scholar

    [39] Cyranoski D. (2019). Artificial intelligence is selecting grant reviewers in China. Nature 569:316−317. DOI:10.1038/d41586-019-01517-8

    View in Article CrossRef Google Scholar

    [40] Skarlinski M.D., Cox S., Laurent J.M., et al. (2024). Language agents achieve superhuman synthesis of scientific knowledge. arXiv preprint. DOI:10.48550/arXiv.2409.13740

    View in Article Google Scholar

    [41] Lin E., Peng Z. and Fang Y. (2025). Evaluating and enhancing large language models for novelty assessment in scholarly publications. Proc. Workshop AI Sci. Discov. 2025:46−57. DOI:10.18653/v1/2025.aisd-main.5

    View in Article CrossRef Google Scholar

    [42] Radensky M., Shahid S., Fok R., et al. (2024). Scideator: Human-llm scientific idea generation grounded in research-paper facet recombination. arXiv preprint. DOI:10.48550/arXiv.2409.14634

    View in Article Google Scholar

    [43] Couto P.H., Ho Q.P., Kumari N., et al. (2024). Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance. arXiv preprint. DOI:10.48550/arXiv.2406.10294

    View in Article Google Scholar

    [44] Faizullah A.R.B.M., Urlana A. and Mishra R. (2024). Limgen: Probing the llms for generating suggestive limitations of research papers. Proc. Jt. Eur. Conf. Mach. Learn. Knowl. Discov. Databases 2024:106−124. DOI:10.1007/978-3-031-70344-7_7

    View in Article CrossRef Google Scholar

    [45] Bhatia C., Pradhan T. and Pal S. (2020). Metagen: An academic meta-review generation system. Proc. Int. ACM SIGIR Conf. Res. Dev. Inf. Retr. 2020:1653−1656. DOI:10.1145/3397271.3401190

    View in Article CrossRef Google Scholar

    [46] Shen C., Cheng L., Zhou R., et al. (2022). MReD: A meta-review dataset for structure-controllable text generation. Findings Assoc. Comput. Linguist. ACL 2022:2521−2535. DOI:10.18653/v1/2022.findings-acl.198

    View in Article CrossRef Google Scholar

    [47] Zeng Q., Sidhu M., Blume A., et al. (2024). Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation. Proc. Int. Jt. Conf. Artif. Intell. 2024:20−38. DOI:10.1007/978-981-97-9536-9_2

    View in Article CrossRef Google Scholar

    [48] Li M., Hovy E. and Lau J. (2023). Summarizing multiple documents with conversational structure for meta-review generation. Findings Assoc. Comput. Linguist. EMNLP 2023:7089−7112. DOI:10.18653/v1/2023.findings-emnlp.472

    View in Article CrossRef Google Scholar

    [49] Sun L., Tao S., Hu J., et al. (2024). MetaWriter: Exploring the the potential and perils of ai writing support in scientific peer review. Proc. ACM Hum.-Comput. Interact. 8:1−32. DOI:10.1145/3637371

    View in Article CrossRef Google Scholar

    [50] Darrin M., Arous I., Piantanida P., et al. (2024). Glimpse: Pragmatically informative multi-document summarization for scholarly reviews. Proc. Annu. Meet. Assoc. Comput. Linguist. 2024:12737−12752. DOI:10.18653/v1/2024.acl-long.688

    View in Article CrossRef Google Scholar

    [51] Sukpanichnant P., Rapberger A. and Toni F. (2024). Peerarg: Argumentative peer review with llms. arXiv preprint. DOI:10.48550/arXiv.2409.16813

    View in Article Google Scholar

    [52] Hossain E., Sinha S.K., Bansal N., et al. (2025). Llms as meta-reviewers’ assistants: A case study. Proc. Conf. North Am. Chapter Assoc. Comput. Linguist. Hum. Lang. Technol. 2025:7763−7803. DOI:10.18653/v1/2025.naacl-long.395

    View in Article CrossRef Google Scholar

    [53] Krizhevsky A., Sutskever I. and Hinton G.E. (2017). ImageNet classification with deep convolutional neural networks. Commun. ACM 60:84−90. DOI:10.1145/3065386

    View in Article CrossRef Google Scholar

    [54] Hinton G., Deng L., Yu D., et al. (2012). Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Process. Mag. 29:82−97. DOI:10.1109/msp.2012.2205597

    View in Article CrossRef Google Scholar

    [55] Devlin J., Chang M.-W., Lee K., et al. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. Proc. NAACL HLT:4171–4186. DOI:10.18653/v1/N19-1423

    View in Article Google Scholar

    [56] Szegedy C., Zaremba W., Sutskever I., et al. (2013). Intriguing properties of neural networks. arXiv preprint. DOI:10.48550/arXiv.1312.6199

    View in Article Google Scholar

    [57] Biggio B. and Roli F. (2018). Wild patterns: Ten years after the rise of adversarial machine learning. Proc. ACM SIGSAC Conf. Comput. Commun. Secur. 2018:2154−2156. DOI:10.1145/3243734.3264418

    View in Article CrossRef Google Scholar

    [58] Goodfellow I.J., Shlens J. and Szegedy C. (2014). Explaining and harnessing adversarial examples. arXiv preprint. DOI:10.48550/arXiv.1412.6572

    View in Article Google Scholar

    [59] Athalye A., Carlini N. and Wagner D. (2018). Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. Proc. Int. Conf. Mach. Learn. 2018:274−283. DOI:10.48550/arXiv.1802.00420

    View in Article CrossRef Google Scholar

    [60] Barreno M., Nelson B., Sears R., et al. (2006). Can machine learning be secure? Proc. ACM Symp. Inf. Comput. Commun. Secur. 2006:16−25. DOI:10.1145/1128817.1128824

    View in Article CrossRef Google Scholar

    [61] Biggio B., Corona I., Maiorca D., et al. (2013). Evasion attacks against machine learning at test time. Proc. Jt. Eur. Conf. Mach. Learn. Knowl. Discov. Databases 2013:387−402. DOI:10.1007/978-3-642-40994-3_25

    View in Article CrossRef Google Scholar

    [62] Carlini N. and Wagner D. (2017). Towards evaluating the robustness of neural networks. Proc. IEEE Symp. Secur. Priv. 2017:39−57. DOI:10.1109/SP.2017.49

    View in Article CrossRef Google Scholar

    [63] Papernot N., McDaniel P., Jha S., et al. (2016). The limitations of deep learning in adversarial settings. Proc. IEEE Eur. Symp. Secur. Priv. 2016:372−387. DOI:10.1109/EuroSP.2016.36

    View in Article CrossRef Google Scholar

    [64] Madry A., Makelov A., Schmidt L., et al. (2017). Towards deep learning models resistant to adversarial attacks. arXiv preprint. DOI:10.48550/arXiv.1706.06083

    View in Article Google Scholar

    [65] Chen P.-Y., Zhang H., Sharma Y., et al. (2017). Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. Proc. ACM Workshop Artif. Intell. Secur. 2017:15−26. DOI:10.1145/3128572.3140448

    View in Article CrossRef Google Scholar

    [66] Ilyas A., Engstrom L., Athalye A., et al. (2018). Black-box adversarial attacks with limited queries and information. Proc. Int. Conf. Mach. Learn. 2018:2137−2146. DOI:10.48550/arXiv.1804.08598

    View in Article CrossRef Google Scholar

    [67] Papernot N., McDaniel P., Sinha A., et al. (2018). Sok: Security and privacy in machine learning. Proc. IEEE Eur. Symp. Secur. Priv. 2018:399−414. DOI:10.1109/EuroSP.2018.00035

    View in Article CrossRef Google Scholar

    [68] Fredrikson M., Jha S. and Ristenpart T. (2015). Model inversion attacks that exploit confidence information and basic countermeasures. Proc. ACM SIGSAC Conf. Comput. Commun. Secur. 2015:1322−1333. DOI:10.1145/2810103.2813677

    View in Article CrossRef Google Scholar

    [69] Shokri R., Stronati M., Song C., et al. (2017). Membership inference attacks against machine learning models. Proc. IEEE Symp. Secur. Priv. 2017:3−18. DOI:10.1109/SP.2017.41

    View in Article CrossRef Google Scholar

    [70] Tramèr F., Zhang F., Juels A., et al. (2016). Stealing machine learning models via prediction APIs. Proc. USENIX Secur. Symp. 2016:601−618. DOI:10.5555/3241094.3241142

    View in Article CrossRef Google Scholar

    [71] Yeom S., Giacomelli I., Fredrikson M., et al. (2018). Privacy risk in machine learning: Analyzing the connection to overfitting. Proc. IEEE Comput. Secur. Found. Symp. 2018:268−282. DOI:10.1109/CSF.2018.00027

    View in Article CrossRef Google Scholar

    [72] Biggio B., Nelson B. and Laskov P. (2012). Poisoning attacks against support vector machines. arXiv preprint. DOI:10.5555/3042573.3042761

    View in Article Google Scholar

    [73] Tolpegin V., Truex S., Gursoy M.E., et al. (2020). Data poisoning attacks against federated learning systems. Proc. Eur. Symp. Res. Comput. Secur. 2020:480−501. DOI:10.1007/978-3-030-58951-6_24

    View in Article CrossRef Google Scholar

    [74] Gu T., Dolan-Gavitt B. and Garg S. (2017). Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint. DOI:10.48550/arXiv.1708.06733

    View in Article Google Scholar

    [75] Chen X., Liu C., Li B., et al. (2017). Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint. DOI:10.48550/arXiv.1712.05526

    View in Article Google Scholar

    [76] Shafahi A., Huang W.R., Najibi M., et al. (2018). Poison frogs! Targeted clean-label poisoning attacks on neural networks. Adv. Neural Inf. Process. Syst. 31:6106−6116. DOI:10.5555/3327345.3327509

    View in Article CrossRef Google Scholar

    [77] Zhang J., Chen B., Cheng X., et al. (2021). PoisonGAN: Generative poisoning attacks against federated learning in edge computing systems. IEEE Internet Things J. 8:3310−3322. DOI:10.1109/jiot.2020.3023126

    View in Article CrossRef Google Scholar

    [78] Carlini N., Athalye A., Papernot N., et al. (2019). On evaluating adversarial robustness. arXiv preprint. DOI:10.48550/arXiv.1902.06705

    View in Article Google Scholar

    [79] Tramèr F., Kurakin A., Papernot N., et al. (2017). Ensemble adversarial training: Attacks and defenses. arXiv preprint. DOI:10.48550/arXiv.1705.07204

    View in Article Google Scholar

    [80] Cohen J., Rosenfeld E. and Kolter Z. (2019). Certified adversarial robustness via randomized smoothing. Proc. Int. Conf. Mach. Learn. 2019:1310−1320. DOI:10.48550/arXiv.1902.02918

    View in Article CrossRef Google Scholar

    [81] Wu D., Xia S.-T. and Wang Y. (2020). Adversarial weight perturbation helps robust generalization. Adv. Neural Inf. Process. Syst. 33:2958−2969. DOI:10.5555/3495724.3495973

    View in Article CrossRef Google Scholar

    [82] Chen T., Liu S., Chang S., et al. (2020). Adversarial robustness: From self-supervised pre-training to fine-tuning. Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. 2020:699−708. DOI:10.1109/CVPR42600.2020.00078

    View in Article CrossRef Google Scholar

    [83] Metzen J.H., Genewein T., Fischer V., et al. (2017). On detecting adversarial perturbations. arXiv preprint. DOI:10.48550/arXiv.1702.04267

    View in Article Google Scholar

    [84] Steinhardt J., Koh P.W.W. and Liang P.S. (2017). Certified defenses for data poisoning attacks. Adv. Neural Inf. Process. Syst. 30:3520−3532. DOI:10.5555/3294996.3295110

    View in Article CrossRef Google Scholar

    [85] Piet J., Alrashed M., Sitawarin C., et al. (2024). Jatmo: Prompt injection defense by task-specific finetuning. Proc. Eur. Symp. Res. Comput. Secur. 2024:105−124. DOI:10.1007/978-3-031-70879-4_6

    View in Article CrossRef Google Scholar

    [86] Doskaliuk B., Zimba O., Yessirkepov M., et al. (2025). Artificial Intelligence in peer review: Enhancing efficiency while preserving integrity. J. Korean Med. Sci. 40:e92. DOI:10.3346/jkms.2025.40.e92https://www.ncbi.nlm.nih.gov/pubmed/39995259

    View in Article CrossRef Google Scholar

    [87] Mann S.P., Aboy M., Seah J.J., et al. (2025). AI and the future of academic peer review. arXiv preprint. DOI:10.48550/arXiv.2509.14189

    View in Article Google Scholar

    [88] Maturo F., Porreca A. and Porreca A. (2025). The risks of artificial intelligence in research: Ethical and methodological challenges in the peer review process. AI Ethics 5:5389−5396. DOI:10.1007/s43681-025-00775-9

    View in Article CrossRef Google Scholar

    [89] Keuper J. (2025). Prompt injection attacks on llm generated reviews of scientific publications. arXiv preprint. DOI:10.48550/arXiv.2509.10248

    View in Article Google Scholar

    [90] Zhao Y., Liu H., Yu D., et al. (2025). One token to fool llm-as-a-judge. arXiv preprint. DOI:10.48550/arXiv.2507.08794

    View in Article Google Scholar

    [91] Collu M.G., Salviati U., Confalonieri R., et al. (2025). Publish to perish: Prompt injection attacks on llm-assisted peer review. arXiv preprint. DOI:10.48550/arXiv.2508.20863

    View in Article Google Scholar

    [92] Ye R., Pang X., Chai J., et al. (2024). Are we there yet? Revealing the risks of utilizing large language models in scholarly peer review. arXiv preprint. DOI:10.48550/arXiv.2412.01708

    View in Article Google Scholar

    [93] Dong Y., Jiang X., Liu H., et al. (2024). Generalization or memorization: Data contamination and trustworthy evaluation for large language models. arXiv preprint. DOI:10.18653/v1/2024.findings-acl.716

    View in Article Google Scholar

    [94] Goldblum M., Tsipras D., Xie C., et al. (2023). Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses. IEEE Trans. Pattern Anal. Mach. Intell. 45:1563−1580. DOI:10.1109/TPAMI.2022.3162397

    View in Article CrossRef Google Scholar

    [95] Borgeaud S., Mensch A., Hoffmann J., et al. (2022). Improving language models by retrieving from trillions of tokens. Proc. Int. Conf. Mach. Learn. 2022:2206−2240. DOI:10.48550/arXiv.2112.04426

    View in Article CrossRef Google Scholar

    [96] Souly A., Rando J., Chapman E., et al. (2025). Poisoning attacks on llms require a near-constant number of poison samples. arXiv preprint. DOI:10.48550/arXiv.2510.07192

    View in Article Google Scholar

    [97] Wen J., Si C., Chen Y.-h., et al. (2025). Predicting empirical ai research outcomes with language models. arXiv preprint. DOI:10.48550/arXiv.2506.00794

    View in Article Google Scholar

    [98] Bereska L. and Gavves E. (2024). Mechanistic interpretability for AI safety-A review. arXiv preprint. DOI:10.48550/arXiv.2404.14082

    View in Article Google Scholar

    [99] Lo L.Y. and Qu H. (2025). How good (or bad) are llms at detecting misleading visualizations? IEEE Trans. Vis. Comput. Graph. 31:1116−1125. DOI:10.1109/TVCG.2024.3456333

    View in Article CrossRef Google Scholar

    [100] Tonglet J., Zimny J., Tuytelaars T., et al. (2025). Is this chart lying to me? Automating the detection of misleading visualizations. arXiv preprint. DOI:10.48550/arXiv.2508.21675

    View in Article Google Scholar

    [101] Gallegos I.O., Rossi R.A., Barrow J., et al. (2024). Bias and fairness in large language models: A survey. Comput. Linguist. 50:1097−1179. DOI:10.1162/coli_a_00524

    View in Article CrossRef Google Scholar

    [102] Liu Y., Deng G., Li Y., et al. (2023). Prompt injection attack against llm-integrated applications. arXiv preprint. DOI:10.48550/arXiv.2306.05499

    View in Article Google Scholar

    [103] Zhou X., Qiang Y., Zade S.Z., et al. (2023). Hijacking large language models via adversarial in-context learning. arXiv preprint. DOI:10.48550/arXiv.2311.09948

    View in Article Google Scholar

    [104] Gong Y., Chen Z., Chen M., et al. (2025). Topic-fliprag: Topic-orientated adversarial opinion manipulation attacks to retrieval-augmented generation models. arXiv preprint. DOI:10.5555/3766078.3766274

    View in Article Google Scholar

    [105] Schwinn L., Dobre D., Günnemann S., et al. (2023). Adversarial attacks and defenses in large language models: Old and new threats. arXiv preprint. DOI:10.48550/arxiv.2310.19737

    View in Article Google Scholar

    [106] Raina V., Liusie A. and Gales M. (2024). Is llm-as-a-judge robust? Investigating universal adversarial attacks on zero-shot llm assessment. arXiv preprint. DOI:10.48550/arXiv.2402.14016

    View in Article Google Scholar

    [107] Guo Y., Guo M., Su J., et al. (2024). Bias in large language models: Origin, evaluation, and mitigation. arXiv preprint. DOI:10.48550/arXiv.2411.10915

    View in Article Google Scholar

    [108] Navigli R., Conia S. and Ross B. (2023). Biases in large language models: Origins, inventory, and discussion. ACM J. Data Inf. Qual. 15:1−21. DOI:10.1145/3597307

    View in Article CrossRef Google Scholar

    [109] Angrist J.D. (2014). The perils of peer effects. Lab. Econ. 30:98−108. DOI:10.1016/j.labeco.2014.05.008

    View in Article CrossRef Google Scholar

    [110] Lewis P., Perez E., Piktus A., et al. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Adv. Neural Inf. Process. Syst. 33:9459−9474. DOI:10.5555/3495724.3496517

    View in Article CrossRef Google Scholar

    [111] Schwarzschild A., Goldblum M., Gupta A., et al. (2021). Just how toxic is data poisoning. a unified benchmark for backdoor and data poisoning attacks. Proc. Int. Conf. Mach. Learn. 2021:9389−9398. DOI:10.48550/arXiv.2006.12557

    View in Article CrossRef Google Scholar

    [112] Touvron H., Lavril T., Izacard G., et al. (2023). Llama: Open and efficient foundation language models. arXiv preprint. DOI:10.48550/arXiv.2302.13971

    View in Article Google Scholar

    [113] Bowen D., Murphy B., Cai W., et al. (2025). Scaling trends for data poisoning in LLMs. Proc. AAAI Conf. Artif. Intell. 39:27206−27214. DOI:10.1609/aaai.v39i26.34929

    View in Article CrossRef Google Scholar

    [114] Liu T., Zhang Y., Feng Z., et al. (2024). Beyond traditional threats: A persistent backdoor attack on federated learning. Proc. AAAI Conf. Artif. Intell. 38:21359−21367. DOI:10.1609/aaai.v38i19.30131

    View in Article CrossRef Google Scholar

    [115] Zhu C., Li Y., Rao B., et al. (2025). SPA: Towards more stealth and persistent backdoor attacks in federated learning. arXiv preprint. DOI:10.48550/arXiv.2506.20931

    View in Article Google Scholar

    [116] Tian Z., Cui L., Liang J., et al. (2022). A comprehensive survey on poisoning attacks and countermeasures in machine learning. ACM Comput. Surv. 55:1−35. DOI:10.1145/3551636

    View in Article CrossRef Google Scholar

    [117] Zhao P., Zhu W., Jiao P., et al. (2025). Data poisoning in deep learning: A survey. arXiv preprint. DOI:10.48550/arXiv.2503.22759

    View in Article Google Scholar

    [118] Muñoz-González L., Biggio B., Demontis A., et al. (2017). Towards poisoning of deep learning algorithms with back-gradient optimization. Proc. ACM Workshop Artif. Intell. Secur. 2017:27−38. DOI:10.1145/3128572.3140451

    View in Article CrossRef Google Scholar

    [119] Nourani M., Roy C., Block J.E., et al. (2021). Anchoring bias affects mental model formation and user reliance in explainable AI systems. Proc. Int. Conf. Intell. User Interfaces 2021:340−350. DOI:10.1145/3397481.3450639

    View in Article CrossRef Google Scholar

    [120] Shi F., Chen X., Misra K., et al. (2023). Large language models can be easily distracted by irrelevant context. Proc. Int. Conf. Mach. Learn. 2023:31210−31227. DOI:10.5555/3618408.3619699

    View in Article CrossRef Google Scholar

    [121] Dougrez-Lewis J., Akhter M.E., Ruggeri F., et al. (2025). Assessing the the reasoning capabilities of llms in the context of evidence-based claim verification. Findings Assoc. Comput. Linguist. ACL 2025:20604−20628. DOI:10.18653/v1/2025.findings-acl.1059

    View in Article CrossRef Google Scholar

    [122] Hong R., Zhang H., Pang X., et al. (2024). A closer look at the self-verification abilities of large language models in logical reasoning. Proc. NAACL HLT 2024:900−925. DOI:10.18653/v1/2024.naacl-long.52

    View in Article CrossRef Google Scholar

    [123] Li J., Li Y., Hu X., et al. (2025). Aspect-guided multi-level perturbation analysis of large language models in automated peer review. arXiv preprint. DOI:10.48550/arXiv.2502.12510

    View in Article Google Scholar

    [124] Liang W., Zhang Y., Cao H., et al. (2024). Can large language models provide useful feedback on research papers? A large-scale empirical analysis. NEJM AI 1:AIoa2400196. DOI:10.1056/AIoa2400196

    View in Article CrossRef Google Scholar

    [125] Zhou Z., Li Z., Zhang J., et al. (2025). Corba: Contagious recursive blocking attacks on multi-agent systems based on large language models. arXiv preprint. DOI:10.48550/arXiv.2502.14529

    View in Article Google Scholar

    [126] Zhu S., Zhang R., An B., et al. (2023). Autodan: Interpretable gradient-based adversarial attacks on large language models. arXiv preprint. DOI:10.48550/arXiv.2310.15140

    View in Article Google Scholar

    [127] Zizzo G., Cornacchia G., Fraser K., et al. (2025). Adversarial prompt evaluation: Systematic benchmarking of guardrails against prompt input attacks on llms. arXiv preprint. DOI:10.48550/arXiv.2502.15427

    View in Article Google Scholar

    [128] Bozdag N.B., Mehri S., Tur G., et al. (2025). Persuade me if you can: A framework for evaluating persuasion effectiveness and susceptibility among large language models. arXiv preprint. DOI:10.48550/arXiv.2503.01829

    View in Article Google Scholar

    [129] Salvi F., Horta Ribeiro M., Gallotti R., et al. (2025). On the conversational persuasiveness of GPT-4. Nat. Hum. Behav. 9:1645−1653. DOI:10.1038/s41562-025-02194-6

    View in Article CrossRef Google Scholar

    [130] Liu Y., Yang K., Liu Y., et al. (2023). The shackles of peer review: Unveiling the flaws in the ivory tower. arXiv preprint. DOI:10.48550/arXiv.2310.05966

    View in Article Google Scholar

    [131] Nisbett R.E. and Wilson T.D. (1977). The halo effect: Evidence for unconscious alteration of judgments. J. Pers. Soc. Psychol. 35:250−256. DOI:10.1037/0022-3514.35.4.250

    View in Article CrossRef Google Scholar

    [132] Zhang J., Zhang H., Deng Z., et al. (2022). Investigating fairness disparities in peer review: A language model enhanced approach. arXiv preprint. DOI:10.48550/arXiv.2211.06398

    View in Article Google Scholar

    [133] Fox C.W., Meyer J. and Aimé E. (2023). Double‐blind peer review affects reviewer ratings and editor decisions at an ecology journal. Funct. Ecol. 37:1144−1157. DOI:10.1111/1365-2435.14259

    View in Article CrossRef Google Scholar

    [134] Sun M., Barry Danfa J. and Teplitskiy M. (2021). Does double‐blind peer review reduce bias. Evidence from a top computer science conference. J. Assoc. Inf. Sci. Technol. 73:811−819. DOI:10.1002/asi.24582

    View in Article CrossRef Google Scholar

    [135] Verharen J.P.H. (2023). ChatGPT identifies gender disparities in scientific peer review. eLife 12:RP90230. DOI:10.7554/eLife.90230

    View in Article CrossRef Google Scholar

    [136] Hosseini M. and Horbach S. (2023). Fighting reviewer fatigue or amplifying bias. Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer review. Res. Integr. Peer Rev. 8:4. DOI:10.1186/s41073-023-00133-5

    View in Article CrossRef Google Scholar

    [137] Soneji A., Kokulu F.B., Rubio-Medrano C., et al. (2022). “Flawed, but like democracy we don’t have a better system”: The Experts’ Insights on the peer review process of evaluating security papers. Proc. IEEE Symp. Secur. Priv. 2022:1845−1862. DOI:10.1109/SP46214.2022.9833581

    View in Article CrossRef Google Scholar

    [138] Schramowski P., Turan C., Andersen N., et al. (2022). Large pre-trained language models contain human-like biases of what is right and wrong to do. Nat. Mach. Intell. 4:258−268. DOI:10.1038/s42256-022-00458-8

    View in Article CrossRef Google Scholar

    [139] Li H., Shan S., Wenger E., et al. (2022). Blacklight: Scalable defense for neural networks against query-based black-box attacks. Proc. USENIX Secur. Symp. 2022:2117–2134. https://www.usenix.org/conference/usenixsecurity22/presentation/li-huiying

    View in Article Google Scholar

    [140] Koo R., Lee M., Raheja V., et al. (2024). Benchmarking cognitive biases in large language models as evaluators. Findings Assoc. Comput. Linguist. ACL 2024:517−545. DOI:10.18653/v1/2024.findings-acl.29

    View in Article CrossRef Google Scholar

    [141] Bartos O.J. and Wehr P. (2002). Using conflict theory (Cambridge University Press). DOI:10.2307/1556608

    View in Article Google Scholar

    [142] Robertson Z. (2023). GPT4 is slightly helpful for peer-review assistance: A pilot study. arXiv preprint. DOI:10.48550/arXiv.2307.05492

    View in Article Google Scholar

  • Cite this article:

    Wang J., Liu Y., Xu H., et al. (2026). When AI reviews science: Can we trust the referee? The Innovation Informatics 2:100030. https://doi.org/10.59717/j.xinn-inform.2026.100030
    Wang J., Liu Y., Xu H., et al. (2026). When AI reviews science: Can we trust the referee? The Innovation Informatics 2:100030. https://doi.org/10.59717/j.xinn-inform.2026.100030

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(7)     Tables(3)

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(9514) PDF downloads(5297)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint