Article Contents
REVIEW   Open Access     Cite

A survey on AI agent security: Reasoning, Acting, and Self-Evolving

    Show all affliationsShow less
More Information
  • DownLoad: Full size image
    1. Artificial intelligence (AI) agents are becoming increasingly autonomous and capable.

      Their growing autonomy introduces new security risks across the entire execution process.

      This survey provides an execution cycle taxonomy of AI agent security challenges and solutions.

      We summarize existing threats, defenses, benchmarks, and future research opportunities.

  • Artificial intelligence (AI) agents are rapidly evolving into autonomous systems capable of reasoning, acting, and continual evolving. Their increasing autonomy enables powerful real-world applications but also introduces security risks throughout the entire agent execution cycle. However, the security challenges arising across different stages of the agent execution cycle have not yet been systematically reviewed. In this survey, we provide a comprehensive review of AI agent security from the perspective of the AI agent execution cycle. Specifically, we establish a unified taxonomy that organizes AI agent execution into three sequential stages: Reasoning, Acting, and Self-Evolving. Each stage is further divided into two interconnected components, resulting in six components in total: Thought, Plan, Skill, Tool, Adaptation, and Evolution. We then systematically review the threats, defense mechanisms, and evaluation benchmarks associated with each component, providing a comprehensive understanding of security challenges throughout the complete execution pipeline. Finally, we discuss current limitations and promising future research directions toward building secure, robust, and trustworthy AI agents. We believe that this survey provides a unified understanding of AI agent security and facilitates future research on the secure deployment of increasingly autonomous AI systems.
  • 加载中
  • [1] Yang A., Li A., Yang B., et al. (2025). Qwen3 technical report. arXiv preprint arXiv: 2505.09388

    View in Article Google Scholar

    [2] Chen Z., Wu J., Wang W., et al. (2024). Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition : 24185–24198.

    View in Article Google Scholar

    [3] Liu A., Feng B., Xue B., et al. (2024). Deepseek-v3 technical report. arXiv preprint arXiv: 2412.19437.

    View in Article Google Scholar

    [4] Achiam J., Adler S., Agarwal S., et al. (2023). Gpt-4 technical report. arXiv preprint arXiv: 2303.08774.

    View in Article Google Scholar

    [5] Team G., Anil R., Borgeaud S., et al. (2023). Gemini: A family of highly capable multimodal models. arXiv preprint arXiv: 2312.11805.

    View in Article Google Scholar

    [6] Yao Y., Duan J., Xu K., et al. (2024). A survey on large language model (llm) security and privacy: The good, the bad, and the ugly. High-Confid. Comput. 4:100211. DOI:10.1016/j.hcc.2024.100211

    View in Article CrossRef Google Scholar

    [7] Alberts I. L., Mercolli L., Pyka T., et al. (2023). Large language models (LLM) and Chat-GPT: what will the impact on nuclear medicine be? Eur. J. Nucl. Med. Mol. Imaging 50:1549−1552. DOI:10.1007/s00259-023-06172-w

    View in Article CrossRef Google Scholar

    [8] Li Y., Huang Y., Yang B., et al. (2024). Snapkv: Llm knows what you are looking for before generation. Adv. Neural Inf. Process. Syst. 37:22947−22970. DOI:10.52202/079017-0722

    View in Article CrossRef Google Scholar

    [9] Wu S., Fei H., Qu L., et al. (2023). Next-gpt: Any-to-any multimodal llm. arXiv preprint arXiv: 2309.05519.

    View in Article Google Scholar

    [10] Wang L., Ma C., Feng X., et al. (2024). A survey on large language model based autonomous agents. Front. Comput. Sci. 18:186345. DOI:10.1007/s11704-024-40231-1

    View in Article CrossRef Google Scholar

    [11] Wooldridge M. (1999). Intelligent agents. Multiagent systems: A modern approach to distributed artificial intelligence 1:27−73.

    View in Article Google Scholar

    [12] Jennings N. R. and Wooldridge M. (1998). Applications of intelligent agents. Agent technology: foundations, applications, and markets (Springer), pp: 3–28.

    View in Article Google Scholar

    [13] Nwana H. S. (1996). Software agents: An overview. Knowl. Eng. Rev. 11:205−244. DOI:10.1017/S026988890000789X

    View in Article CrossRef Google Scholar

    [14] Zhao A., Huang D., Xu Q., et al. (2024). Expel: Llm agents are experiential learners. Proceedings of the AAAI conference on artificial intelligence 38:19632−19642.

    View in Article Google Scholar

    [15] Zhou S., Xu F. F., Zhu H., et al. (2024). Webarena: A realistic web environment for building autonomous agents. International conference on learning representations 2024:15585−15606.

    View in Article Google Scholar

    [16] Deng X., Gu Y., Zheng B., et al. (2023). Mind2web: Towards a generalist agent for the web. Adv. Neural Inf. Process. Syst. 36:28091−28114.

    View in Article Google Scholar

    [17] Koh J. Y., Lo R., Jang L., et al. (2024). Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers): 881–905.

    View in Article Google Scholar

    [18] He H., Yao W., Ma K., et al. (2024). Webvoyager: Building an end-to-end web agent with large multimodal models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers): 6864–6890.

    View in Article Google Scholar

    [19] Ning L., Liang Z., Jiang Z., et al. (2025). A survey of webagents: Towards next-generation ai agents for web automation with large foundation models. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2 : 6140–6150.

    View in Article Google Scholar

    [20] Chae H., Kim N., Ong K., et al. (2025). Web agents with world models: Learning and leveraging environment dynamics in web navigation. International conference on learning representations 2025:63707−63738.

    View in Article Google Scholar

    [21] Nguyen D., Chen J., Wang Y., et al. (2025). Gui agents: A survey. Findings of the association for computational linguistics: ACL 2025 : 22522–22538.

    View in Article Google Scholar

    [22] Yao S., Chen H., Yang J., et al. (2022). Webshop: Towards scalable real-world web interaction with grounded language agents. Adv. Neural Inf. Process. Syst. 35:20744−20757.

    View in Article Google Scholar

    [23] Hong W., Wang W., Lv Q., et al. (2024). Cogagent: A visual language model for gui agents. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition : 14281–14290.

    View in Article Google Scholar

    [24] Cheng K., Sun Q., Chu Y., et al. (2024). Seeclick: Harnessing gui grounding for advanced visual gui agents. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) : 9313–9332.

    View in Article Google Scholar

    [25] Gou B., Wang D. R., Zheng B., et al. (2025). Navigating the digital world as humans do: Universal visual grounding for gui agents. International conference on learning representations 2025:30851−30883.

    View in Article Google Scholar

    [26] Zhang C., He S., Qian J., et al. (2024). Large language model-brained gui agents: A survey. arXiv preprint arXiv: 2411.18279.

    View in Article Google Scholar

    [27] Wang S., Liu W., Chen J., et al. (2024). Gui agents with foundation models: A comprehensive survey. arXiv preprint arXiv: 2411.04890.

    View in Article Google Scholar

    [28] Chen W., Cui J., Hu J., et al. (2025). Guicourse: From general vision language model to versatile gui agent. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) : 21936–21959.

    View in Article Google Scholar

    [29] Wu Q., Cheng K., Yang R., et al. (2026). Gui-actor: Coordinate-free visual grounding for gui agents. Adv. Neural Inf. Process. Syst. 38:15101−15128.

    View in Article Google Scholar

    [30] Liu Y., Li P., Wei Z., et al. (2026). Infiguiagent: A multimodal generalist gui agent with native reasoning and reflection. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) : 1035–1051.

    View in Article Google Scholar

    [31] Wang X., Wang B., Lu D., et al. (2026). Opencua: Open foundations for computer-use agents. Adv. Neural Inf. Process. Syst. 38:139756−139806.

    View in Article Google Scholar

    [32] Sager P. J., Meyer B., Yan P., et al. (2026). A comprehensive survey of agents for computer use: Foundations, challenges, and future directions. J. Artif. Intell. Res. 85.

    View in Article Google Scholar

    [33] Agashe S., Wong K., Tu V., et al. (2025). Agent s2: A compositional generalist-specialist framework for computer use agents. arXiv preprint arXiv: 2504.00906.

    View in Article Google Scholar

    [34] Sun Z., Liu Z., Zang Y., et al. (2025). Seagent: Self-evolving computer use agent with autonomous learning from experience. arXiv preprint arXiv: 2508.04700.

    View in Article Google Scholar

    [35] Lai H., Liu X., Zhao Y., et al. (2025). Computerrl: Scaling end-to-end online reinforcement learning for computer use agents. arXiv preprint arXiv: 2508.14040.

    View in Article Google Scholar

    [36] Kuntz T., Duzan A., Zhao H., et al. (2026). Os-harm: A benchmark for measuring safety of computer use agents. Adv. Neural Inf. Process. Syst. 38.

    View in Article Google Scholar

    [37] Li M., Zhao S., Wang Q., et al. (2024). Embodied agent interface: Benchmarking llms for embodied decision making. Adv. Neural Inf. Process. Syst. 37:100428−100534.

    View in Article Google Scholar

    [38] Feng Z., Xue R., Yuan L., et al. (2026). Multi-agent embodied ai: Advances and future directions. Sci. China Inf. Sci. 69:151202. DOI:10.1007/s11432-025-4820-4

    View in Article CrossRef Google Scholar

    [39] Zhang H., Du W., Shan J., et al. (2024). Building cooperative embodied agents modularly with large language models. International conference on learning representations 2024:19373−19401.

    View in Article Google Scholar

    [40] Fan L., Wang G., Jiang Y., et al. (2022). Minedojo: Building open-ended embodied agents with internet-scale knowledge. Adv. Neural Inf. Process. Syst. 35:18343−18362.

    View in Article Google Scholar

    [41] Xia F., Zamir A. R., He Z., et al. (2018). Gibson env: Real-world perception for embodied agents. Proceedings of the IEEE conference on computer vision and pattern recognition : 9068–9079.

    View in Article Google Scholar

    [42] Franklin S. (1997). Autonomous agents as embodied AI. Cybernetics & Systems 28:499−520.

    View in Article Google Scholar

    [43] Dorri A., Kanhere S. S. and Jurdak R. (2018). Multi-agent systems: A survey. IEEE Access 6:28573−28593. DOI:10.1109/ACCESS.2018.2831228

    View in Article CrossRef Google Scholar

    [44] Maldonado D., Cruz E., Torres J. A., et al. (2024). Multi-agent systems: A survey about its components, framework and workflow. IEEE Access 12:80950−80975. DOI:10.1109/ACCESS.2024.3409051

    View in Article CrossRef Google Scholar

    [45] Du H., Thudumu S., Vasa R., et al. (2024). A survey on context-aware multi-agent systems: techniques, challenges and future directions. arXiv preprint arXiv: 2402.01968.

    View in Article Google Scholar

    [46] Li A., Xie Y., Li S., et al. (2025). Agent-oriented planning in multi-agent systems. International conference on learning representations 2025:19495−19517.

    View in Article Google Scholar

    [47] Wang Y. and Chen X. (2025). Mirix: Multi-agent memory system for llm-based agents. arXiv preprint arXiv: 2507.07957.

    View in Article Google Scholar

    [48] Alzubi S., Provenzano N., Bingham J., et al. (2026). Evoskill: Automated skill discovery for multi-agent systems. arXiv preprint arXiv: 2603.02766.

    View in Article Google Scholar

    [49] Qiao S., Fang R., Zhang N., et al. (2024). Agent planning with world knowledge model. Adv. Neural Inf. Process. Syst. 37:114843−114871.

    View in Article Google Scholar

    [50] Xu W., Liang Z., Mei K., et al. (2026). A-mem: Agentic memory for llm agents. Adv. Neural Inf. Process. Syst. 38:17577−17604.

    View in Article Google Scholar

    [51] Kang J., Ji M., Zhao Z., et al. (2025). Memory os of ai agent. Proceedings of the 2025 conference on empirical methods in natural language processing : 25972–25981.

    View in Article Google Scholar

    [52] He Z., Wang Y., Zhi C., et al. (2026). Memoryarena: Benchmarking agent memory in inter-dependent multi-session agentic tasks. arXiv preprint arXiv: 2602.16313.

    View in Article Google Scholar

    [53] Masterman T., Besen S., Sawtell M., et al. (2024). The landscape of emerging ai agent architectures for reasoning, planning, and tool calling: A survey. arXiv preprint arXiv: 2404.11584.

    View in Article Google Scholar

    [54] Rahimi A. (2026). Agentic and multi-agent systems: A systematic review of tool use, benchmarks, and governance.

    View in Article Google Scholar

    [55] Lumer E., Gulati A., Nizar F., et al. (2026). Tool and agent selection for large language model agents in production: A survey. 2026 IEEE conference on artificial intelligence (CAI) : 701–708.

    View in Article Google Scholar

    [56] Liu M. M., Garcia D., Parllaku F., et al. (2026). Toolscope: Enhancing llm agent tool use through tool merging and context-aware filtering. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) : 34095–34119.

    View in Article Google Scholar

    [57] Xu R. and Yan Y. (2026). Agent skills for large language models: Architecture, acquisition, security, and the path forward. arXiv preprint arXiv: 2602.12430.

    View in Article Google Scholar

    [58] Li X., Chen W., Liu Y., et al. (2026). SkillsBench: Benchmarking how well agent skills work across diverse tasks. arXiv preprint arXiv: 2602.12670.

    View in Article Google Scholar

    [59] Ni J., Liu Y., Liu X., et al. (2026). Trace2skill: Distill trajectory-local lessons into transferable agent skills. arXiv preprint arXiv: 2603.25158.

    View in Article Google Scholar

    [60] Yue L., Bhandari K. R., Ko C. Y., et al. (2026). From static templates to dynamic runtime graphs: a survey of workflow optimization for llm agents. arXiv preprint arXiv: 2603.22386.

    View in Article Google Scholar

    [61] Jenkins A., Kitkowska A., Maidhof C., et al. (2026). Security and privacy in agentic AI: grand challenges and future directions. arXiv preprint arXiv: 2607.06608.

    View in Article Google Scholar

    [62] Zhang D., Feng G., Shi Y., et al. (2021). Physical safety and cyber security analysis of multi-agent systems: A survey of recent advances. IEEE/CAA J. Autom. Sin. 8:319−333. DOI:10.1109/JAS.2021.1003820

    View in Article CrossRef Google Scholar

    [63] Cui T., Wang Y., Fu C., et al. (2024). Risk taxonomy, mitigation, and assessment benchmarks of large language model systems. arXiv preprint arXiv: 2401.05778.

    View in Article Google Scholar

    [64] Gan Y., Yang Y., Ma Z., et al. (2024). Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents. CoRR abs/2411.09523 (2024). arXiv preprint arXiv: 2411.09523.

    View in Article Google Scholar

    [65] Liu D., Yang M., Qu X., et al. (2025). A survey of attacks on large vision–language models: Resources, advances, and future trends. IEEE Trans. Neural Netw. Learn. Syst.

    View in Article Google Scholar

    [66] Deng Z., Guo Y., Han C., et al. (2025). Ai agents under threat: A survey of key security challenges and future pathways. ACM Comput. Surv. 57:1−36.

    View in Article Google Scholar

    [67] Luo J., Zhang W., Yuan Y., et al. (2025). Large language model agent: A survey on methodology, applications and challenges. arXiv preprint arXiv: 2503.21460.

    View in Article Google Scholar

    [68] Chen A., Wu Y., Zhang J., et al. (2025). A survey on the safety and security threats of computer-using agents: JARVIS or ultron? arXiv preprint arXiv: 2505.10924.

    View in Article Google Scholar

    [69] Wang K., Zhang G., Zhou Z., et al. (2025). A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment. arXiv preprint arXiv: 2504.15585.

    View in Article Google Scholar

    [70] Ma X., Gao Y., Wang Y., et al. (2026). Safety at scale: A comprehensive survey of large model and agent safety. Found. Trends Priv. Secur. 8:1−240.

    View in Article Google Scholar

    [71] Zhou M., Duan N., Liu S., et al. (2020). Progress in neural NLP: modeling, learning, and reasoning. Engineering 6:275−290. DOI:10.1016/j.eng.2019.12.014

    View in Article CrossRef Google Scholar

    [72] Lopez M. M. and Kalita J. (2017). Deep Learning applied to NLP. arXiv preprint arXiv: 1703.03091.

    View in Article Google Scholar

    [73] Su Y., Lan T., Li H., et al. (2023). Pandagpt: One model to instruction-follow them all. Proceedings of the 1st Workshop on Taming Large Language Models: Controllability in the era of Interactive Assistants! : 11–23.

    View in Article Google Scholar

    [74] Zhou J., Lu T., Mishra S., et al. (2023). Instruction-following evaluation for large language models. arXiv preprint arXiv: 2311.07911.

    View in Article Google Scholar

    [75] Lou R., Zhang K. and Yin W. (2024). Large language model instruction following: A survey of progresses and challenges. Comput. Linguist. 50:1053−1095. DOI:10.1162/coli_a_00523

    View in Article CrossRef Google Scholar

    [76] Wu X., Yao W., Chen J., et al. (2024). From language modeling to instruction following: Understanding the behavior shift in llms after instruction tuning. Proceedings of the 2024 conference of the north american chapter of the association for computational linguistics: Human language technologies (volume 1: Long papers) : 2341–2369.

    View in Article Google Scholar

    [77] Qin Y., Song K., Hu Y., et al. (2024). Infobench: Evaluating instruction following ability in large language models. Findings of the association for computational linguistics: ACL 2024 : 13025–13048.

    View in Article Google Scholar

    [78] Eigner E. and Händler T. (2024). Determinants of llm-assisted decision-making. arXiv preprint arXiv: 2402.17385.

    View in Article Google Scholar

    [79] Ma S., Chen Q., Wang X., et al. (2025). Towards human-ai deliberation: Design and evaluation of llm-empowered deliberative ai for ai-assisted decision-making. Proceedings of the 2025 CHI conference on human factors in computing systems : 1–23.

    View in Article Google Scholar

    [80] Xu F. F., Song Y., Li B., et al. (2026). Theagentcompany: benchmarking llm agents on consequential real world tasks. Adv. Neural Inf. Process. Syst. 38.

    View in Article Google Scholar

    [81] Dong G., Lu J., Huang J., et al. (2026). Agent-world: Scaling real-world environment synthesis for evolving general agent intelligence. arXiv preprint arXiv: 2604.18292.

    View in Article Google Scholar

    [82] Sun S., Song H., Huang L., et al. (2026). SWE-world: Building software engineering agents in docker-free environments. arXiv preprint arXiv: 2602.03419.

    View in Article Google Scholar

    [83] Ferrag M. A., Tihanyi N. and Debbah M. (2026). From llm reasoning to autonomous ai agents: A comprehensive review. IEEE Access.

    View in Article Google Scholar

    [84] Putta P., Mills E., Garg N., et al. (2024). Agent q: Advanced reasoning and learning for autonomous ai agents. arXiv preprint arXiv: 2408.07199.

    View in Article Google Scholar

    [85] Zhou A., Yan K., Shlapentokh-Rothman M., et al. (2023). Language agent tree search unifies reasoning acting and planning in language models. arXiv preprint arXiv: 2310.04406.

    View in Article Google Scholar

    [86] Fang J., Peng Y., Zhang X., et al. (2025). A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems. arXiv preprint arXiv: 2508.07407.

    View in Article Google Scholar

    [87] Gupta A., Savarese S., Ganguli S., et al. (2021). Embodied intelligence via learning and evolution. Nat. Commun. 12:5721. DOI:10.1038/s41467-021-25874-z

    View in Article CrossRef Google Scholar

    [88] Zhai Y., Tao S., Chen C., et al. (2025). Agentevolver: Towards efficient self-evolving agent system. arXiv preprint arXiv: 2511.10395.

    View in Article Google Scholar

    [89] Shao S., Ren Q., Qian C., et al. (2025). Your agent may misevolve: Emergent risks in self-evolving llm agents. arXiv preprint arXiv: 2509.26354.

    View in Article Google Scholar

    [90] Liu Y., Zhang R., Luo H., et al. (2025). Secure multi-LLM agentic AI and agentification for edge general intelligence by zero-trust: A survey. arXiv preprint arXiv: 2508.19870.

    View in Article Google Scholar

    [91] Zhou L., Varadharajan V. and Hitchens M. (2013). Achieving secure role-based access control on encrypted data in cloud storage. IEEE Trans. Inf. Forensics Secur. 8:1947−1960. DOI:10.1109/TIFS.2013.2286456

    View in Article CrossRef Google Scholar

    [92] Huang K., Mehmood Y., Atta H., et al. (2025). Fortifying the Agentic Web: A Unified Zero-Trust Architecture Against Logic-layer Threats. arXiv preprint arXiv: 2508.12259.

    View in Article Google Scholar

    [93] Andreotta A. J., Kirkham N. and Rizzi M. (2022). AI, big data, and the future of consent. AI Soc 37:1715−1728. DOI:10.1007/s00146-021-01262-5

    View in Article CrossRef Google Scholar

    [94] Lindqvist H. (2006). Mandatory access control. Master’s thesis in computing science, Umea University, Department of Computing Science, SE-901 87.

    View in Article Google Scholar

    [95] Wei J., Wang X., Schuurmans D., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst. 35:24824−24837. DOI:10.59350/hqv3w-d3q61

    View in Article CrossRef Google Scholar

    [96] Zhang Z., Zhang A., Li M., et al. (2022). Automatic chain of thought prompting in large language models. arXiv preprint arXiv: 2210.03493.

    View in Article Google Scholar

    [97] Lyu Q., Havaldar S., Stein A., et al. (2023). Faithful chain-of-thought reasoning. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) : 305–329.

    View in Article Google Scholar

    [98] Chen Q., Qin L., Liu J., et al. (2026). Towards reasoning era: A survey of long chain-of-thought for reasoning large language models. Sci. China Inf. Sci. 69:161101. DOI:10.1007/s11432-025-4665-8

    View in Article CrossRef Google Scholar

    [99] Feng G., Zhang B., Gu Y., et al. (2023). Towards revealing the mystery behind chain of thought: a theoretical perspective. Adv. Neural Inf. Process. Syst. 36:70757−70798.

    View in Article Google Scholar

    [100] Hu Y., Cai Y., Du Y., et al. (2025). Self-evolving multi-agent collaboration networks for software development. International conference on learning representations 2025:23007−23039.

    View in Article Google Scholar

    [101] Ou Y., Zhou W., Ding S., et al. (2025). Symbolic learning enables self-evolving agents. AI Open 6:314−322. DOI:10.1016/j.aiopen.2025.11.004

    View in Article CrossRef Google Scholar

    [102] Ouyang S., Yan J., Hsu I., et al. (2025). Reasoningbank: Scaling agent self-evolving with reasoning memory. arXiv preprint arXiv: 2509.25140.

    View in Article Google Scholar

    [103] Wang X., Wei J., Schuurmans D., et al. (2022). Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv: 2203.11171.

    View in Article Google Scholar

    [104] Tan X., Wang X., Liu Q., et al. (2025). Paths-over-graph: Knowledge graph empowered large language model reasoning. Proceedings of the ACM on web conference 2025:3505−3522.

    View in Article Google Scholar

    [105] Li W., Wei W., Qu X., et al. (2023). Trea: Tree-structure reasoning schema for conversational recommendation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) : 2970–2982.

    View in Article Google Scholar

    [106] Jin B., Xie C., Zhang J., et al. (2024). Graph chain-of-thought: Augmenting large language models by reasoning on graphs. Findings of the association for computational linguistics: ACL 2024 : 163–184.

    View in Article Google Scholar

    [107] Kuo M., Zhang J., Ding A., et al. (2025). H-cot: Hijacking the chain-of-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek-r1, and gemini 2.0 flash thinking. arXiv preprint arXiv: 2502.12893.

    View in Article Google Scholar

    [108] Ye C., Cui J. and Hadfield-Menell D. (2026). Prompt injection as role confusion. arXiv preprint arXiv: 2603.12277.

    View in Article Google Scholar

    [109] Xiang Z., Jiang F., Xiong Z., et al. (2024). Badchain: Backdoor chain-of-thought prompting for large language models. arXiv preprint arXiv: 2401.12242.

    View in Article Google Scholar

    [110] Liu S., Li R., Yu L., et al. (2026). Badthink: Triggered overthinking attacks on chain-of-thought reasoning in large language models. Proceedings of the AAAI conference on artificial intelligence 40:32141−32149.

    View in Article Google Scholar

    [111] Cao Y., Gu N., Shen X., et al. (2024). Defending large language models against jailbreak attacks through chain of thought prompting. 2024 International conference on networking and network applications (NaNA) : 125–130.

    View in Article Google Scholar

    [112] Wang W., Hosseini P. and Feizi S. (2025). Chain-of-defensive-thought: Structured reasoning elicits robustness in large language models against reference corruption. arXiv preprint arXiv: 2504.20769.

    View in Article Google Scholar

    [113] Xue Z., Bi Z., Ma L., et al. (2025). Thought purity: A defense framework for chain-of-thought attack. arXiv preprint arXiv: 2507.12314.

    View in Article Google Scholar

    [114] Jacovi A., Bitton Y., Bohnet B., et al. (2024). A chain-of-thought is as strong as its weakest link: A benchmark for verifiers of reasoning chains. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) : 4615–4634.

    View in Article Google Scholar

    [115] Jiang D., Zhang R., Guo Z., et al. (2025). Mme-cot: Benchmarking chain-of-thought in large multimodal models for reasoning quality, robustness, and efficiency. arXiv preprint arXiv: 2502.09621.

    View in Article Google Scholar

    [116] Chen Q., Qin L., Zhang J., et al. (2024). M3cot: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) : 8199–8221.

    View in Article Google Scholar

    [117] Zhang H., Huang J., Mei K., et al. (2025). Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents. International conference on learning representations 2025:35331−35366.

    View in Article Google Scholar

    [118] Li Z., Chang Y. and Wu Y. (2025). Think-bench: Evaluating thinking efficiency and chain-of-thought quality of large reasoning models. arXiv preprint arXiv: 2505.22113.

    View in Article Google Scholar

    [119] Bailey L., Ong E., Russell S., et al. (2023). Image hijacks: Adversarial images can control generative models at runtime. arXiv preprint arXiv: 2309.00236.

    View in Article Google Scholar

    [120] Cui X., Aparcedo A., Jang Y. K., et al. (2024). On the robustness of large multimodal models against image adversarial attacks. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition : 24625–24634.

    View in Article Google Scholar

    [121] Fu X., Wang Z., Li S., et al. (2023). Misusing tools in large language models with visual adversarial examples. arXiv preprint arXiv: 2310.03185.

    View in Article Google Scholar

    [122] Gao K., Bai Y., Gu J., et al. (2024). Inducing high energy-latency of large vision-language models with verbose images. arXiv preprint arXiv: 2401.11170.

    View in Article Google Scholar

    [123] Luo H., Gu J., Liu F., et al. (2024). An image is worth 1000 lies: Adversarial transferability across prompts on vision-language models. arXiv preprint arXiv: 2403.09766.

    View in Article Google Scholar

    [124] Rahmatullaev T., Druzhinina P., Kurdiukov N., et al. (2025). Universal adversarial attack on aligned multimodal llms. arXiv preprint arXiv: 2502.07987.

    View in Article Google Scholar

    [125] Schlarmann C. and Hein M. (2023). On the adversarial robustness of multi-modal foundation models. Proceedings of the IEEE/CVF international conference on computer vision : 3677–3685.

    View in Article Google Scholar

    [126] Shao Z., Liu H., Hu Y., et al. (2024). Refusing safe prompts for multi-modal large language models. arXiv preprint arXiv: 2407.09050.

    View in Article Google Scholar

    [127] Tan Z., Zhao C., Moraffah R., et al. (2024). The wolf within: Covert injection of malice into mllm societies via an mllm operative. arXiv preprint arXiv: 2402.14859.

    View in Article Google Scholar

    [128] Wang Z., Han Z., Chen S., et al. (2024). Stop reasoning! when multimodal llm with chain-of-thought reasoning meets adversarial image. arXiv preprint arXiv: 2402.14899.

    View in Article Google Scholar

    [129] Wu C. H., Shah R., Koh J. Y., et al. (2024). Dissecting adversarial robustness of multimodal LM agents. arXiv preprint arXiv: 2406.12814.

    View in Article Google Scholar

    [130] Dong Y., Chen H., Chen J., et al. (2023). How robust is google’s bard to adversarial image attacks? arXiv preprint arXiv: 2309.11751.

    View in Article Google Scholar

    [131] Guo Q., Pang S., Jia X., et al. (2024). Efficient generation of targeted and transferable adversarial examples for vision-language models via diffusion models. IEEE Trans. Inf. Forensics Secur. 20:1333−1348.

    View in Article Google Scholar

    [132] Kim H. S., Kim M. and Kim C. (2024). Doubly-universal adversarial perturbations: Deceiving vision-language models across both images and text with a single perturbation. arXiv preprint arXiv: 2412.08108.

    View in Article Google Scholar

    [133] Tu H., Cui C., Wang Z., et al. (2024). How many are in this image a safety evaluation benchmark for vision LLMs. European conference on computer vision : 37–55.

    View in Article Google Scholar

    [134] Wang H., Dong K., Zhu Z., et al. (2024). Transferable multimodal attack on vision-language pre-training models. 2024 IEEE symposium on security and privacy (SP) : 1722–1740.

    View in Article Google Scholar

    [135] Wang X., Ji Z., Ma P., et al. (2023). Instructta: Instruction-tuned targeted attack for large vision-language models. arXiv preprint arXiv: 2312.01886.

    View in Article Google Scholar

    [136] Wang Y., Liu C., Qu Y., et al. (2024). Break the visual perception: Adversarial attacks targeting encoded visual tokens of large vision-language models. Proceedings of the 32nd ACM International Conference on Multimedia : 1072–1081.

    View in Article Google Scholar

    [137] Zhang J., Ye J., Ma X., et al. (2025). Anyattack: Towards large-scale self-supervised adversarial attacks on vision-language models. Proceedings of the computer vision and pattern recognition conference : 19900–19909.

    View in Article Google Scholar

    [138] Zhao Y., Pang T., Du C., et al. (2023). On evaluating adversarial robustness of large vision-language models. Adv. Neural Inf. Process. Syst. 36:54111−54138.

    View in Article Google Scholar

    [139] Liu D., Yang M., Qu X., et al. (2024). Pandora’s box: Towards building universal attackers against real-world large vision-language models. Adv. Neural Inf. Process. Syst. 37:52127−52158.

    View in Article Google Scholar

    [140] Yang Y., Wang L., Yang X., et al. (2025). Effective black-box multi-faceted attacks breach vision large language model guardrails. arXiv preprint arXiv: 2502.05772.

    View in Article Google Scholar

    [141] Zhang H., Shao W., Liu H., et al. (2024). B-avibench: Toward evaluating the robustness of large vision-language model on black-box adversarial visual-instructions. IEEE Trans. Inf. Forensics Secur. 20:1434−1446.

    View in Article Google Scholar

    [142] Carlini N., Nasr M., Choquette-Choo C. A., et al. (2023). Are aligned neural networks adversarially aligned? Adv. Neural Inf. Process. Syst. 36:61478−61500.

    View in Article Google Scholar

    [143] Chen T., Wang K. and Wei H. (2026). Zer0-Jack: A memory-efficient gradient-based jail-breaking method for black box Multi-modal Large Language Models. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) : 4328–4344.

    View in Article Google Scholar

    [144] Gu X., Zheng X., Pang T., et al. (2024). Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast. arXiv preprint arXiv: 2402.08567.

    View in Article Google Scholar

    [145] Zaitang L., Chen P. Y. and Ho T. Y. (2025). Retention score: Quantifying jailbreak risks for vision language models. Proceedings of the AAAI conference on artificial intelligence 39:27446−27454.

    View in Article Google Scholar

    [146] Niu Z., Ren H., Gao X., et al. (2024). Jailbreaking attack against multimodal large language model. arXiv preprint arXiv: 2402.02309.

    View in Article Google Scholar

    [147] Oh S., Jin Y., Sharma M., et al. (2024). Uniguard: Towards universal safety guardrails for jailbreak attacks on multimodal large language models. arXiv preprint arXiv: 2411.01703.

    View in Article Google Scholar

    [148] Schaeffer R., Valentine D., Bailey L., et al. (2024). When do universal image jailbreaks transfer between vision-language models? Workshop on responsibly building the next generation of multimodal foundational models.

    View in Article Google Scholar

    [149] Wang R., Ma X., Zhou H., et al. (2024). White-box multimodal jailbreaks against large vision-language models. Proceedings of the 32nd ACM International Conference on Multimedia : 6920–6928.

    View in Article Google Scholar

    [150] Gong Y., Ran D., Liu J., et al. (2025). Figstep: Jailbreaking large vision-language models via typographic visual prompts. Proceedings of the AAAI conference on artificial intelligence 39:23951−23959.

    View in Article Google Scholar

    [151] Li Y., Guo H., Zhou K., et al. (2024). Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models. European conference on computer vision : 174–189.

    View in Article Google Scholar

    [152] Liu X., Zhu Y., Gu J., et al. (2024). Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. European conference on computer vision : 386–403.

    View in Article Google Scholar

    [153] Liu Y., Cai C., Zhang X., et al. (2024). Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts. Proceedings of the 32nd ACM International Conference on Multimedia : 3578–3586.

    View in Article Google Scholar

    [154] Luo W., Ma S., Liu X., et al. (2024). Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks. arXiv preprint arXiv: 2404.03027.

    View in Article Google Scholar

    [155] Ma S., Luo W., Wang Y., et al. (2024). Visual-roleplay: Universal jailbreak attack on multimodal large language models via role-playing image character. arXiv preprint arXiv: 2405.20773.

    View in Article Google Scholar

    [156] Teng M., Xiaojun J., Ranjie D., et al. (2024). Heuristic-induced multi-modal risk distribution jailbreak attack for multimodal large language models. arXiv preprint arXiv: 2412.05934.

    View in Article Google Scholar

    [157] Wang Y., Zhou X., Wang Y., et al. (2025). Jailbreak large vision-language models through multi-modal linkage. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) : 1466–1494.

    View in Article Google Scholar

    [158] Wu Y., Li X., Liu Y., et al. (2023). Jailbreaking gpt-4v via self-adversarial attacks with system prompts. arXiv preprint arXiv: 2311.09127.

    View in Article Google Scholar

    [159] Zhao S., Duan R., Wang F., et al. (2025). Jailbreaking multimodal large language models via shuffle inconsistency. Proceedings of the IEEE/CVF international conference on computer vision : 2045–2054.

    View in Article Google Scholar

    [160] Wei R., Niu P., Shen X., et al. (2025). The trojan knowledge: Bypassing commercial llm guardrails via harmless prompt weaving and adaptive tree search. arXiv preprint arXiv: 2512.01353.

    View in Article Google Scholar

    [161] Xiong C., Chen P. Y. and Ho T. Y. (2026). CoP: agentic red-teaming for large language models using composition of principles. Adv. Neural Inf. Process. Syst. 38:104257−104291.

    View in Article Google Scholar

    [162] Bagdasaryan E., Hsieh T. Y., Nassi B., et al. (2023). Abusing images and sounds for indirect instruction injection in multi-modal LLMs. arXiv preprint arXiv: 2307.10490.

    View in Article Google Scholar

    [163] Kimura S., Tanaka R., Miyawaki S., et al. (2024). Empirical analysis of large vision-language models against goal hijacking via visual prompt injection. arXiv preprint arXiv: 2408.03554.

    View in Article Google Scholar

    [164] Qraitem M., Tasnim N., Teterwak P., et al. (2024). Vision-llms can fool themselves with self-generated typographic attacks. arXiv preprint arXiv: 2402.00626.

    View in Article Google Scholar

    [165] Chen Y., Mendes E., Das S., et al. (2023). Can language models be instructed to protect personal information? arXiv preprint arXiv: 2310.02224.

    View in Article Google Scholar

    [166] Qi X., Huang K., Panda A., et al. (2024). Visual adversarial examples jailbreak aligned large language models. Proceedings of the AAAI conference on artificial intelligence 38:21527−21536.

    View in Article Google Scholar

    [167] Shayegani E., Dong Y. and Abu-Ghazaleh N. (2024). Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models. International conference on learning representations 2024:30853−30885.

    View in Article Google Scholar

    [168] Yang Y., Li Y., Yao H., et al. (2026). ShadowCode: Towards (automatic) external prompt injection attack against code LLMs. IEEE Trans. Dependable Secure Comput. 23:10527-10540. DOI:10.1109/TDSC.2026.3703498

    View in Article CrossRef Google Scholar

    [169] Xu Y., Yao J., Shu M., et al. (2024). Shadowcast: Stealthy data poisoning attacks against vision-language models. Adv. Neural Inf. Process. Syst. 37:57733−57764.

    View in Article Google Scholar

    [170] Liang J., Liang S., Liu A., et al. (2025). Vl-trojan: Multimodal instruction backdoor attacks against autoregressive visual language models. Int. J. Comput. Vis. 133:3994−4013. DOI:10.1007/s11263-025-02368-9

    View in Article CrossRef Google Scholar

    [171] Liang S., Liang J., Pang T., et al. (2025). Revisiting backdoor attacks against large vision-language models from domain shift. Proceedings of the computer vision and pattern recognition conference : 9477–9486.

    View in Article Google Scholar

    [172] Lu D., Pang T., Du C., et al. (2024). Test-time backdoor attacks on multimodal large language models. arXiv preprint arXiv: 2402.08577.

    View in Article Google Scholar

    [173] Ni Z., Ye R., Wei Y., et al. (2024). Physical backdoor attack can jeopardize driving with vision-large-language models. arXiv preprint arXiv: 2404.12916.

    View in Article Google Scholar

    [174] Tao X., Zhong S., Li L., et al. (2025). Imgtrojan: Jailbreaking vision-language models with one image. Proceedings of the 2025 conference of the nations of the americas chapter of the association for computational linguistics: Human language technologies (volume 1: Long papers) : 7048–7063.

    View in Article Google Scholar

    [175] Xiong C., Qi X., Chen P. Y., et al. (2025). Defensive prompt patch: A robust and generalizable defense of large language models against jailbreak attacks. Findings of the association for computational linguistics: ACL 2025 : 409–437.

    View in Article Google Scholar

    [176] Hu X., Chen P. Y. and Ho T. Y. (2025). Token highlighter: Inspecting and mitigating jailbreak prompts for large language models. Proceedings of the AAAI conference on artificial intelligence 39:27330−27338.

    View in Article Google Scholar

    [177] Hu X., Chen P. Y. and Ho T. Y. (2024). Gradient cuff: Detecting jailbreak attacks on large language models by exploring refusal loss landscapes. Adv. Neural Inf. Process. Syst. 37:126265−126296.

    View in Article Google Scholar

    [178] Hung K. H., Ko C. Y., Rawat A., et al. (2025). Attention tracker: Detecting prompt injection attacks in llms. Findings of the association for computational linguistics: NAACL 2025 : 2309–2322.

    View in Article Google Scholar

    [179] Chern S., Fan Z. and Liu A. (2024). Combating adversarial attacks with multi-agent debate. arXiv preprint arXiv: 2401.05998.

    View in Article Google Scholar

    [180] Lin G., Tanaka T. and Zhao Q. (2024). Large language model sentinel: Llm agent for adversarial purification. arXiv preprint arXiv: 2405.20770.

    View in Article Google Scholar

    [181] Barua S., Rahman M., Sadek M. J., et al. (2025). Guardians of the agentic system: Preventing many shots jailbreak with agentic system. arXiv preprint arXiv: 2502.16750.

    View in Article Google Scholar

    [182] Zeng Y., Wu Y., Zhang X., et al. (2024). Autodefense: Multi-agent llm defense against jailbreak attacks. arXiv preprint arXiv: 2403.04783.

    View in Article Google Scholar

    [183] Ni Z., Wang H. and Wang H. (2025). Shieldlearner: A new paradigm for jailbreak attack defense in llms. arXiv preprint arXiv: 2502.13162.

    View in Article Google Scholar

    [184] Chang M., Chhablani G., Clegg A., et al. (2025). Partnr: A benchmark for planning and reasoning in embodied multi-agent tasks. International conference on learning representations 2025:65205−65268.

    View in Article Google Scholar

    [185] Valmeekam K., Marquez M., Olmo A., et al. (2023). Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change. Adv. Neural Inf. Process. Syst. 36:38975−38987.

    View in Article Google Scholar

    [186] Geng L. and Chang E. Y. (2025). Realm-bench: A real-world planning benchmark for llms and multi-agent systems. arXiv preprint arXiv: 2502.18836.

    View in Article Google Scholar

    [187] Yin S., Pang X., Ding Y., et al. (2024). Safeagentbench: A benchmark for safe task planning of embodied llm agents. arXiv preprint arXiv: 2412.13178.

    View in Article Google Scholar

    [188] Wang Z., Li Y., Wu Y., et al. (2026). Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents. arXiv preprint arXiv: 2606.13385.

    View in Article Google Scholar

    [189] Schmotz D., Abdelnabi S. and Andriushchenko M. (2025). Agent skills enable a new class of realistic and trivially simple prompt injections. arXiv preprint arXiv: 2510.26328.

    View in Article Google Scholar

    [190] Liu Y., Wang W., Feng R., et al. (2026). Agent skills in the wild: An empirical study of security vulnerabilities at scale. arXiv preprint arXiv: 2601.10338.

    View in Article Google Scholar

    [191] Liu Y., Chen Z., Zhang Y., et al. (2026). Malicious agent skills in the wild: A large-scale security empirical study. arXiv preprint arXiv: 2602.06547.

    View in Article Google Scholar

    [192] Liu S., Li C., Wang C., et al. (2026). Clawkeeper: Comprehensive safety protection for openclaw agents through skills, plugins, and watchers. arXiv preprint arXiv: 2603.24414.

    View in Article Google Scholar

    [193] Schmotz D., Beurer-Kellner L., Abdelnabi S., et al. (2026). Skill-inject: Measuring agent vulnerability to skill file attacks. arXiv preprint arXiv: 2602.20156.

    View in Article Google Scholar

    [194] Jin C., Wang A., Wei Z., et al. (2026). SkillSafetyBench: Evaluating agent safety under skill-facing attack surfaces. arXiv preprint arXiv: 2605.12015.

    View in Article Google Scholar

    [195] Jiang Y., Zhang Y., Backes M., et al. (2026). HarmfulSkillBench: How do harmful skills weaponize your agents? arXiv preprint arXiv: 2604.15415.

    View in Article Google Scholar

    [196] Betser R., Bose S., Giloni A., et al. (2026). AgenTRIM: Tool risk mitigation for agentic AI. arXiv preprint arXiv: 2601.12449.

    View in Article Google Scholar

    [197] Hu Y., Jia Y., Li M., et al. (2026). MalTool: Malicious tool attacks on LLM agents. arXiv preprint arXiv: 2602.12194.

    View in Article Google Scholar

    [198] Zhan Q., Liang Z., Ying Z., et al. (2024). Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents. Findings of the association for computational linguistics: ACL 2024 : 10471–10506.

    View in Article Google Scholar

    [199] Lee L. F., Chang Y. Y., Yu C. M., et al. (2026). WebMCP tool surface poisoning: Runtime manipulation attacks on LLM agents. arXiv preprint arXiv: 2606.06387.

    View in Article Google Scholar

    [200] Liu L., Han T., Liu Z., et al. (2026). ShareLock: A stealthy multi-tool threshold poisoning attack against MCP. arXiv preprint arXiv: 2606.27027.

    View in Article Google Scholar

    [201] He Y., Zhu H., Li Y., et al. (2026). AttriGuard: Defeating indirect prompt injection in LLM agents via causal attribution of tool invocations. arXiv preprint arXiv: 2603.10749.

    View in Article Google Scholar

    [202] Ye H., Zhang Z., Jia J., et al. (2026). Trustdesc: Preventing tool poisoning in llm applications via trusted description generation. arXiv preprint arXiv: 2604.07536.

    View in Article Google Scholar

    [203] Zhao W., Li Z., Zhang P., et al. (2026). ClawGuard: A runtime security framework for tool-augmented LLM agents against indirect prompt injection. arXiv preprint arXiv: 2604.11790.

    View in Article Google Scholar

    [204] Sigdel A. and Baral R. (2026). ToolMisuseBench: An offline deterministic benchmark for tool misuse and recovery in agentic systems. arXiv preprint arXiv: 2604.01508.

    View in Article Google Scholar

    [205] Mou Y., Xue Z., Li L., et al. (2026). ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback. arXiv preprint arXiv: 2601.10156.

    View in Article Google Scholar

    [206] Wang Z., Gao Y., Wang Y., et al. (2026). Mcptox: A benchmark for tool poisoning on real-world mcp servers. Proceedings of the AAAI conference on artificial intelligence 40:35811−35819.

    View in Article Google Scholar

    [207] Li X., Yu S., Pan M., et al. (2026). Unsafer in many turns: Benchmarking and defending multi-turn safety risks in tool-using agents. arXiv preprint arXiv: 2602.13379.

    View in Article Google Scholar

    [208] Madaan A., Tandon N., Gupta P., et al. (2023). Self-refine: Iterative refinement with selffeedback. Adv. Neural Inf. Process. Syst. 36:46534−46594.

    View in Article Google Scholar

    [209] Zelikman E., Wu Y. and Goodman N. D. (2022). Star: Self-taught reasoner. Proceedings of the NIPS 22:3.

    View in Article Google Scholar

    [210] Hosseini A., Yuan X., Malkin N., et al. V-star: Training verifiers for self-taught reasoners, 2024. URL https://arxiv.org/abs/2402.06457.

    View in Article Google Scholar

    [211] Weng Y., Zhu M., Xia F., et al. (2023). Large language models are better reasoners with self-verification. Findings of the association for computational linguistics: EMNLP 2023 : 2550–2575.

    View in Article Google Scholar

    [212] Wang K., Lou J., Zhou Z., et al. (2026). OEP: Poisoning self-evolving LLM agents via locally correct but non-transferable experiences. arXiv preprint arXiv: 2605.18930.

    View in Article Google Scholar

    [213] Wang Z., Wang K., Wang Q., et al. (2025). Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning. arXiv preprint arXiv: 2504.20073.

    View in Article Google Scholar

    [214] Yu T., Yang Y., Luo X., et al. (2026). UNSEEN: A cross-stack LLM unlearning defense against AR-LLM social engineering attacks. arXiv preprint arXiv: 2604.23141.

    View in Article Google Scholar

    [215] Zheng J., Xu J., Luo Y., et al. (2026). Self-Guard: Defending Large Reasoning Models via enhanced self-reflection. arXiv preprint arXiv: 2602.00707.

    View in Article Google Scholar

    [216] Wang H., Wang Y., Li H., et al. (2026). Be your own red teamer: Safety alignment via self-play and reflective experience replay. arXiv preprint arXiv: 2601.10589.

    View in Article Google Scholar

    [217] Zhong Q., Ding L., Liu J., et al. (2023). Self-evolution learning for discriminative language model pretraining. Findings of the association for computational linguistics: ACL 2023 : 4130–4145.

    View in Article Google Scholar

    [218] Wu S., Lu K., Xu B., et al. (2023). Self-evolved diverse data sampling for efficient instruction tuning. arXiv preprint arXiv: 2311.08182.

    View in Article Google Scholar

    [219] Yuan W., Pang R. Y., Cho K., et al. (2024). Self-rewarding language models. arXiv preprint arXiv: 2401.10020.

    View in Article Google Scholar

    [220] Yang K., Klein D., Celikyilmaz A., et al. (2024). RLCD: Reinforcement learning from contrastive distillation for LM alignment. International conference on learning representations 2024:21179−21218.

    View in Article Google Scholar

    [221] Pang J. C., Wang P., Li K., et al. (2023). Language model self-improvement by reinforcement learning contemplation. arXiv preprint arXiv: 2305.14483.

    View in Article Google Scholar

    [222] Wu R., Wang X., Mei J., et al. (2025). Evolver: Self-evolving llm agents through an experience-driven lifecycle. arXiv preprint arXiv: 2510.16079.

    View in Article Google Scholar

    [223] Zhao A., Wu Y., Wu T., et al. (2026). Absolute zero: Reinforced self-play reasoning with zero data. Adv. Neural Inf. Process. Syst. 38:105816−105879.

    View in Article Google Scholar

    [224] Han S., Xiong K., Liu J., et al. (2025). Alignment tipping process: How self-evolution pushes llm agents off the rails. arXiv preprint arXiv: 2510.04860.

    View in Article Google Scholar

    [225] Huang T., Hu S., Ilhan F., et al.Safety tax: Safety alignment makes your large reasoning models less reasonable, 2025. URL https://arxiv.org/abs/2503.00555.

    View in Article Google Scholar

    [226] Li J. and Kim J. E. (2025). Safety alignment can be not superficial with explicit safety signals. arXiv preprint arXiv: 2505.17072.

    View in Article Google Scholar

    [227] Li J., Zhang Z., Zhou S., et al. (2026). When safe models merge into danger: Exploiting latent vulnerabilities in LLM fusion. arXiv preprint arXiv: 2604.00627.

    View in Article Google Scholar

    [228] Zhang D., Liu X., Cheng L., et al. (2026). Selaur: Self evolving llm agent via uncertainty-aware rewards. Pacific-asia conference on knowledge discovery and data mining : 424–436.

    View in Article Google Scholar

    [229] Huang J., Cheng F., Jiang J., et al. (2026). BenchTrace: A benchmark for testing reflection ability and controlled evolution in LLM agents. arXiv preprint arXiv: 2605.29225.

    View in Article Google Scholar

    [230] Jiang S., Ma L., Hong Z., et al. (2026). SEA-eval: A benchmark for evaluating self-evolving agents beyond episodic assessment. arXiv preprint arXiv: 2604.08988.

    View in Article Google Scholar

    [231] Chan A., Salganik R., Markelius A., et al. (2023). Harms from increasingly agentic algorithmic systems. Proceedings of the 2023 ACM conference on fairness, accountability, and transparency : 651–666.

    View in Article Google Scholar

    [232] Vallor S. and Vierkant T. (2024). Find the gap: AI, responsible agency and vulnerability. Minds Mach. (Dordr.) 34:20. DOI:10.1007/s11023-024-09674-0

    View in Article CrossRef Google Scholar

    [233] European Parliament and Council of the European Union (2016). Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation).

    View in Article Google Scholar

    [234] European Parliament and Council of the European Union (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 on laying down harmonised rules on Artificial Intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (Artificial Intelligence Act).

    View in Article Google Scholar

    [235] Baldoni M., Baroglio C., Micalizio R., et al. (2023). Accountability in multi-agent organizations: from conceptual design to agent programming. Auton. Agent Multi-Agent Syst. 37:7. DOI:10.1007/s10458-022-09590-6

    View in Article CrossRef Google Scholar

    [236] Acharya D. B., Kuppan K. and Divya B. (2025). Agentic AI: Autonomous intelligence for complex goals—A comprehensive survey. IEEE Access 13:18912−18936. DOI:10.1109/ACCESS.2025.3532853

    View in Article CrossRef Google Scholar

    [237] Kapoor S., Gruver N., Roberts M., et al. (2024). Large language models must be taught to know what they don’t know. Adv. Neural Inf. Process. Syst. 37:85932−85972.

    View in Article Google Scholar

    [238] Susser D., Roessler B. and Nissenbaum H. (2019). Online manipulation: Hidden influences in a digital world. Geo. L. Tech. Rev. 4:1.

    View in Article Google Scholar

    [239] Guidotti R. (2024). Counterfactual explanations and how to find them: literature review and benchmarking: R. guidotti. Data Min. Knowl. Discov. 38:2770−2824. DOI:10.1007/s10618-022-00831-6

    View in Article CrossRef Google Scholar

    [240] Giovannoni C., Metta C., Monreale A., et al. (2026). A survey on multimodal explainable Artificial Intelligence. Intell. Syst. Appl. : 200671.

    View in Article Google Scholar

    [241] OpenAI (2026). Sandbox agents.

    View in Article Google Scholar

    [242] High-Level Expert Group on Artificial Intelligence (2019). Ethics guidelines for trustworthy AI. (European Commission).

    View in Article Google Scholar

    [243] OECD (2019). Recommendation of the council on artificial intelligence.

    View in Article Google Scholar

    [244] UNESCO (2021). Recommendation on the ethics of artificial intelligence.

    View in Article Google Scholar

    [245] The IEEE Global Initiative on Ethics of Autonomous and Intelligent Systems (2019). Ethically Aligned Design: A Vision for Prioritizing Human Well-being with Autonomous and Intelligent Systems: First Edition. (IEEE).

    View in Article Google Scholar

    [246] Floridi L. and Cowls J. (2019). A unified framework of five principles for AI in society. Harv. Data Sci. Rev. 1. DOI:10.1162/99608f92.8cd550d1

    View in Article CrossRef Google Scholar

    [247] Aristotle (2000). Nicomachean ethics. Crisp R. (ed). (Cambridge University Press).

    View in Article Google Scholar

    [248] Kant I. (2012). Groundwork of the metaphysics of morals. Timmermann J. (ed). (Cambridge University Press).

    View in Article Google Scholar

    [249] Mill J. S. (2001). Utilitarianism. Sher G. (ed). (Hackett Publishing Company).

    View in Article Google Scholar

    [250] Habermas J. (1990). Moral consciousness and communicative action. (MIT Press).

    View in Article Google Scholar

  • Cite this article:

    Liu D., Luo T., Huang S., et al. (2026). A survey on AI agent security: Reasoning, Acting, and Self-Evolving. AI Plus 1:100004. https://doi.org/10.59717/ipj.aiplus.2026.100004
    Liu D., Luo T., Huang S., et al. (2026). A survey on AI agent security: Reasoning, Acting, and Self-Evolving. AI Plus 1:100004. https://doi.org/10.59717/ipj.aiplus.2026.100004

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(5)     Tables(5)

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(111) PDF downloads(15)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint