Multimodal information-fused embodied intelligence toward autonomous robotic microsurgery
Autonomous surgery offers potential to enhance consistency, improve patient safety, reduce differences in surgeon performance, shorten training periods, and decrease dependency on human resources. Autonomy levels for medical robotics are no autonomy, robot assistance, task autonomy, conditional autonomy, high autonomy, and full autonomy. Currently, task and conditional autonomy have been achieved in laparoscopic and intracavitary surgery. Nevertheless, high-precision microsurgery still relies on the surgeon's teleoperation for sensing, decision-making, and execution, with little related work on autonomous systems. At the microscale, minute motion or force errors may cause irreversible tissue injury; therefore, autonomy is not a matter of simple threshold tightening but requires tightly coupled spatial perception, instrument motion awareness, and real-time safe control with microsurgery-specific validation. Embodied intelligence systems capable of perception, autonomous navigation, and real-world interaction integrate artificial intelligence into surgical robots through tight body-environment coupling, offering a promising solution for achieving high autonomy and full autonomy. This commentary examines how multimodal information-fused embodied intelligence motivates autonomous robotic ophthalmological surgery, which is one of the most delicate and scale-sensitive microsurgical procedures.
Multimodal image fusion for augmented spatial perception of the global scene
High autonomy in robotic microsurgery requires global spatial perception of the entire intraocular scene, covering both the anterior and posterior segments, together with local task cues that typically focus on the retina-contact region. Imaging modalities involve inherent conflicts, where higher resolution narrows the range, faster data acquisition limits the coverage, and susceptibility to noise reduces reliability. For example, the microscope’s fixed viewpoint prevents capturing structures beneath the retinal surface, while image quality suffers from occlusion, uneven lighting, and lens artifacts. Fusing static wide-field, dynamic, and tomographic imagery could enhance spatial perception.
Embodied intelligence in multiple imaging modalities enables geometric or semantic understanding of anatomical structures, enhancing spatial perception in robotic surgery. Preoperative anatomical information is obtained through modalities such as fundus photographs (FPs), fluorescein angiography (FA), wide-field scanning laser ophthalmoscopy (SLO), and optical coherence tomography (OCT). Intraoperative OCT (iOCT) provides high-resolution tomographic retinal images and monitors anatomical changes during surgery. These images offer a more comprehensive view of ocular tissue structure than the microscope, with robust multimodal models already used for retinal lesion monitoring, diagnosis, and localization.
