Article Contents
ARTICLE   Open Access     Cite

SafeMLRM: Demystifying safety in multi-modal large reasoning models

More Information
  • Corresponding authors: xiangwang@ustc.edu.cn (X.W.);  hexn@ustc.edu.cn (X.H.)
  • DownLoad: Full size image
    1. Multi-modal large reasoning models (MLRMs) become less safe after reasoning training.

      Safety failures vary sharply by topic, revealing severe and scenario-specific blind spots.

      Unsafe reasoning usually reaches the final answer, but limited self-correction can still occur.

      OpenSafeMLRM offers a unified toolkit for evaluating models, datasets, and jailbreak attacks.

  • The rapid advancement of multi-modal large reasoning models (MLRMs) — enhanced versions of multi-modal large language models (MLLMs) equipped with reasoning capabilities — has revolutionized diverse applications. However, their safety implications remain underexplored. While prior work has exposed critical vulnerabilities in unimodal reasoning models, MLRMs introduce distinct risks from cross-modal reasoning pathways. This work presents the first systematic safety analysis of MLRMs through large-scale empirical studies comparing MLRMs with their base MLLMs. Our experiments reveal three critical findings: (1) The Reasoning Tax: Acquiring reasoning capabilities catastrophically degrades inherited safety alignment. MLRMs exhibit 37.44% higher jailbreaking success rates than base MLLMs under adversarial attacks. (2) Safety Blind Spots: While safety degradation is pervasive, certain scenarios (e.g., Illegal Activity) suffer 25×higher attack rates — far exceeding the average 3.4×increase, revealing scenario-specific vulnerabilities with alarming cross-model and datasets consistency. (3) Emergent Self-Correction: Despite tight reasoning-answer safety coupling, MLRMs demonstrate nascent self-correction — 16.9% of jailbroken reasoning steps are overridden by safe answers, hinting at intrinsic safeguards. These findings underscore the urgency of scenario-aware safety auditing and mechanisms to amplify MLRMs’ self-correction potential. To catalyze research, we open-source OpenSafeMLRM, the first toolkit for MLRM safety evaluation, providing unified interface for mainstream models, datasets, and jailbreaking methods. Our work calls for immediate efforts to harden reasoning-augmented AI, ensuring its transformative potential aligns with ethical safeguards.
  • 加载中
  • [1] Guo D., Yang D., Zhang H. et al. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645:633−638. DOI:10.1038/s41586-025-09422-z

    View in Article CrossRef Google Scholar

    [2] OpenAI (2025). Introducing OpenAI o3 and o4-mini. OpenAI, April 16, 2025. https://openai.com/index/introducing-o3-and-o4-mini

    View in Article Google Scholar

    [3] Yang S., Tong Y., Niu X. et al. (2025). Demystifying long chain-of-thought reasoning. Proc. Mach. Learn. Res. 267:71177–71209. https://proceedings.mlr.press/v267/yang25ae.html.

    View in Article Google Scholar

    [4] Guo D., Zhu Q., Yang D. et al. (2024). DeepSeek-Coder: When the large language model meets programming—The rise of code intelligence. arXiv preprint arXiv:2401.14196. DOI:10.48550/arXiv.2401.14196.

    View in Article Google Scholar

    [5] Shao Z., Wang P., Zhu Q. et al. (2024). DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. DOI:10.48550/arXiv.2402.03300.

    View in Article Google Scholar

    [6] Zhang J., Huang J., Yao H. et al. (2025). R1-VL: Learning to reason with multimodal large language models via Step-wise group relative policy optimization. Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV):1859–1869. DOI:10.1109/ICCV51701.2025.00181.

    View in Article Google Scholar

    [7] Zhou H., Li X., Wang R. et al. (2025). R1-Zero's “Aha Moment” in visual reasoning on a 2B non-SFT model. arXiv preprint arXiv:2503.05132. DOI:10.48550/arXiv.2503.05132.

    View in Article Google Scholar

    [8] Zhang R., Zhang B., Li Y. et al. (2025). Improve vision language model chain-of-thought reasoning. Proc. 63rd Annu. Meet. Assoc. Comput. Linguist. (ACL), 1:1631–1662. DOI:10.18653/v1/2025.acl-long.82.

    View in Article Google Scholar

    [9] Shen H., Liu P., Li J. et al. (2025). VLM-R1: A stable and generalizable R1-style large vision-language model. arXiv preprint arXiv:2504.07615. DOI:10.48550/arXiv.2504.07615.

    View in Article Google Scholar

    [10] Liu Y., Peng B., Zhong Z. et al. (2025). Seg-Zero: Reasoning-chain guided segmentation via cognitive reinforcement. arXiv preprint arXiv:2503.06520. DOI:10.48550/arXiv.2503.06520.

    View in Article Google Scholar

    [11] Dong Y., Liu Z., Sun H.L. et al. (2025). Insight-V: Exploring long-chain visual reasoning with multimodal large language models. Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR):9062–9072. DOI:10.1109/CVPR52734.2025.00847.

    View in Article Google Scholar

    [12] EvolvingLMMs-Lab (2025). Multimodal Open R1. GitHub repository. https://github.com/EvolvingLMMs-Lab/open-r1-multimodal

    View in Article Google Scholar

    [13] Meng F., Du L., Liu Z. et al. (2025). MM-Eureka: Exploring visual Aha Moment with rule-based large-scale reinforcement learning. arXiv preprint arXiv:2503.07365. DOI:10.48550/arXiv.2503.07365.

    View in Article Google Scholar

    [14] Chen Z., Luo X. and Li D. (2025). VisRL: Intention-driven visual perception via reinforced reasoning. Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV):2545–2555. DOI:10.1109/ICCV51701.2025.00245.

    View in Article Google Scholar

    [15] Zhao J., Wei X. and Bo L. (2025). R1-Omni: Explainable omni-multimodal emotion recognition with reinforcement learning. arXiv preprint arXiv:2503.05379. DOI:10.48550/arXiv.2503.05379.

    View in Article Google Scholar

    [16] Himakunthala V., Ouyang A., Rose D. et al. (2023). Let's think frame by frame with VIP: A video infilling and prediction dataset for evaluating video chain-of-thought. Proc. 2023 Conf. Empir. Methods Nat. Lang. Process. (EMNLP):204–219. DOI:10.18653/v1/2023.emnlp-main.15.

    View in Article Google Scholar

    [17] Meng F., Yang H., Wang Y. et al. (2023). Chain of images for intuitively reasoning. arXiv preprint arXiv:2311.09241. DOI:10.48550/arXiv.2311.09241.

    View in Article Google Scholar

    [18] Xie J., Lei S., Yu Y. et al. (2025). Leveraging chain of thought towards empathetic spoken dialogue without corresponding question-answering data. Proc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP):1–5. DOI:10.1109/ICASSP49660.2025.10889870.

    View in Article Google Scholar

    [19] Luo X., Ding F., Song Y. et al. (2026). PKRD-CoT: A unified chain-of-thought prompting for multi-modal large language models in autonomous driving. Mahmud M., Doborjeh M., Wong K. et al. (eds). Neural Information Processing: 31st International Conference, ICONIP 2024, Proceedings, Part XIII (Springer), Commun. Comput. Inf. Sci. 2294, pp: 62–76. DOI:10.1007/978-981-96-7008-6_5.

    View in Article Google Scholar

    [20] Zheng H., Xu T., Sun H. et al. (2024). Thinking before looking: Improving multimodal LLM reasoning via mitigating visual hallucination. arXiv preprint arXiv:2411.12591. DOI:10.48550/arXiv.2411.12591.

    View in Article Google Scholar

    [21] Gao T., Chen P., Zhang M. et al. (2024). Cantor: Inspiring multimodal chain-of-thought of MLLM. Proc. 32nd ACM Int. Conf. Multimedia:9096–9105. DOI:10.1145/3664647.3681249.

    View in Article Google Scholar

    [22] Wu W., Mao S., Zhang Y. et al. (2024). Mind's eye of LLMs: Visualization-of-thought elicits spatial reasoning in large language models. Adv. Neural Inf. Process. Syst. 37:90277−90317. DOI:10.52202/079017-2866

    View in Article CrossRef Google Scholar

    [23] Luan B., Feng H., Chen H. et al. (2026). TextCoT: Zoom-In for enhanced multimodal text-rich image understanding. ACM Trans. Multim. Comput. Commun. Appl. 22:106:1–106:19. DOI:10.1145/3785474.

    View in Article Google Scholar

    [24] Wang Y., Wu S., Zhang Y. et al. (2025). Multimodal chain-of-thought reasoning: A comprehensive survey. arXiv preprint arXiv:2503.12605. DOI:10.48550/arXiv.2503.12605.

    View in Article Google Scholar

    [25] Battersby S. (2025). Gen AI Reasoning Model Comparison & AI Safety Results. Chatterbox, January 27, 2025. https://chatterbox.co/blog/gen-ai-reasoning-model-comparison-ai-safety-results

    View in Article Google Scholar

    [26] Singh S. (2025). Introducing Safety Aligned DeepSeek R1 Model by Enkrypt AI. Enkrypt AI, January 31, 2025. https://www.enkryptai.com/blog/introducing-safety-aligned-deepseek-r1-model-by-enkrypt-ai

    View in Article Google Scholar

    [27] Zhang W., Lei X., Liu Z. et al. (2025). Safety evaluation of DeepSeek models in Chinese contexts. arXiv preprint arXiv:2502.11137. DOI:10.48550/arXiv.2502.11137.

    View in Article Google Scholar

    [28] Xu Z., Gardiner J. and Belguith S. (2025). The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models. arXiv preprint arXiv:2502.01225. DOI:10.48550/arXiv.2502.01225.

    View in Article Google Scholar

    [29] Wang H., Qin Z., Shen L. et al. (2025). Safety reasoning with guidelines. Proc. Mach. Learn. Res. 267:64084–64108. https://proceedings.mlr.press/v267/wang25cg.html.

    View in Article Google Scholar

    [30] Parmar M. and Govindarajulu Y. (2025). Challenges in ensuring AI safety in DeepSeek-R1 models: The shortcomings of reinforcement learning strategies. arXiv preprint arXiv:2501.17030. DOI:10.48550/arXiv.2501.17030.

    View in Article Google Scholar

    [31] Kumar A., Roh J., Naseh A. et al. (2025). OverThink: Slowdown attacks on reasoning LLMs. arXiv preprint arXiv:2502.02542. DOI:10.48550/arXiv.2502.02542.

    View in Article Google Scholar

    [32] Chen Q., Qin L., Wang J. et al. (2024). Unlocking the capabilities of thought: A reasoning boundary framework to quantify and optimize chain-of-thought. Adv. Neural Inf. Process. Syst. 37:54872−54904. DOI:10.52202/079017-1740

    View in Article CrossRef Google Scholar

    [33] Ying Z., Zheng G., Huang Y. et al. (2025). Towards understanding the safety boundaries of DeepSeek models: Evaluation and findings. arXiv preprint arXiv:2503.15092. DOI:10.48550/arXiv.2503.15092.

    View in Article Google Scholar

    [34] Huang T., Hu S., Ilhan F. et al. (2025). Safety tax: Safety alignment makes your large reasoning models less reasonable. arXiv preprint arXiv:2503.00555. DOI:10.48550/arXiv.2503.00555.

    View in Article Google Scholar

    [35] Ye M., Rong X., Huang W. et al. (2025). A survey of safety on large vision-language models: Attacks, defenses and evaluations. arXiv preprint arXiv:2502.14881. DOI:10.48550/arXiv.2502.14881.

    View in Article Google Scholar

    [36] Liu X., Zhu Y., Gu J. et al. (2024). MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models. Lect. Notes Comput. Sci. 15114:386−403. DOI:10.1007/978-3-031-72992-8_22

    View in Article CrossRef Google Scholar

    [37] Gong Y., Ran D., Liu J. et al. (2025). FigStep: Jailbreaking large vision-language models via typographic visual prompts. Proc. AAAI Conf. Artif. Intell. 39:23951−23959. DOI:10.1609/aaai.v39i22.34568

    View in Article CrossRef Google Scholar

    [38] OpenAI, Achiam J., Adler S. et al. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774. DOI:10.48550/arXiv.2303.08774.

    View in Article Google Scholar

    [39] Rombach R., Blattmann A., Lorenz D. et al. (2022). High-resolution image synthesis with latent diffusion models. Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR):10684–10695. DOI:10.1109/CVPR52688.2022.01042.

    View in Article Google Scholar

    [40] Touvron H., Martin L., Stone K. et al. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. DOI:10.48550/arXiv.2307.09288.

    View in Article Google Scholar

    [41] Yang Y., He X., Pan H. et al. (2025). R1-Onevision: Advancing generalized multimodal reasoning through cross-modal formalization. Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV):2376–2385. DOI:10.1109/ICCV51701.2025.00229.

    View in Article Google Scholar

    [42] Yao H., Huang J., Wu W. et al. (2025). Mulberry: Empowering MLLM with o1-like reasoning and reflection via collective Monte Carlo tree search. Adv. Neural Inf. Process. Syst. 38:89159−89192. DOI:10.52202/085713-1004

    View in Article CrossRef Google Scholar

    [43] Bai S., Chen K., Liu X. et al. (2025). Qwen2.5-VL technical report. arXiv preprint arXiv:2502.13923. DOI:10.48550/arXiv.2502.13923.

    View in Article Google Scholar

    [44] Wang P., Bai S., Tan S. et al. (2024). Qwen2-VL: Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191. DOI:10.48550/arXiv.2409.12191.

    View in Article Google Scholar

    [45] Li B., Zhang K., Zhang H. et al. (2024). LLaVA-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild. LLaVA-VL Blog, May 10, 2024. https://llava-vl.github.io/blog/2024-05-10-llava-next-stronger-llms

    View in Article Google Scholar

    [46] Zheng Y., Lu J., Wang S. et al. (2025). EasyR1: An efficient, scalable, multi-modality RL training framework. GitHub repository. https://github.com/hiyouga/EasyR1

    View in Article Google Scholar

    [47] Yang Y., He X., Pan H. et al. (2025). R1-Onevision: Open-source multimodal large language model with reasoning ability. Project page, February 13, 2025. https://yangyi-vai.notion.site/r1-onevision

    View in Article Google Scholar

    [48] Chen L., Li L., Zhao H. et al. (2025). R1-V: Reinforcing super generalization ability in vision-language models with less than $3. GitHub repository. https://github.com/StarsfieldAI/R1-V

    View in Article Google Scholar

    [49] Peng Y., Zhang G., Zhang M. et al. (2025). LMM-R1: Empowering 3B LMMs with strong reasoning abilities through two-stage rule-based RL. arXiv preprint arXiv:2503.07536. DOI:10.48550/arXiv.2503.07536.

    View in Article Google Scholar

    [50] Arrieta A., Ugarte M., Valle P. et al. (2025). o3-mini vs DeepSeek-R1: Which one is safer? arXiv preprint arXiv:2501.18438. DOI:10.48550/arXiv.2501.18438.

    View in Article Google Scholar

    [51] Mazeika M., Phan L., Yin X. et al. (2024). HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. Proc. Mach. Learn. Res. 235:35181–35224. https://proceedings.mlr.press/v235/mazeika24a.html.

    View in Article Google Scholar

    [52] Zhou K., Liu C., Zhao X. et al. (2025). The hidden risks of large reasoning models: A safety assessment of R1. Proc. 14th Int. Joint Conf. Nat. Lang. Process. and 4th Conf. Asia-Pacific Chapter Assoc. Comput. Linguist. (IJCNLP-AACL):3250–3265. DOI:10.18653/v1/2025.ijcnlp-long.173.

    View in Article Google Scholar

    [53] Lou X., Li Y., Xu J. et al. (2025). Think in safety: Unveiling and mitigating safety alignment collapse in multimodal large reasoning model. Proc. 2025 Conf. Empir. Methods Nat. Lang. Process. (EMNLP):5167–5186. DOI:10.18653/v1/2025.emnlp-main.261.

    View in Article Google Scholar

  • Cite this article:

    Fang J., Wang Y., Wang R., et al. (2026). SafeMLRM: Demystifying safety in multi-modal large reasoning models. AI Plus 1:100014. https://doi.org/10.59717/ipj.aiplus.2026.100014
    Fang J., Wang Y., Wang R., et al. (2026). SafeMLRM: Demystifying safety in multi-modal large reasoning models. AI Plus 1:100014. https://doi.org/10.59717/ipj.aiplus.2026.100014

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(2)     Tables(2)

Supplementary Information

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(35) PDF downloads(9)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint