Artificial intelligence facilitates information fusion for perception in complex environments

COMMENTARY Open Access Download: PDF

Multimodal perception is a foundational technology for human perception in complex environments. These environments often involve various interference conditions and sensor technical limitations that constrain the information capture capabilities of single-modality sensors. Multimodal perception addresses these by integrating complementary multisource heterogeneous information, providing a solution for perceiving complex environments. This technology spans across fields such as autonomous driving, industrial detection, biomedical engineering, and remote sensing. However, challenges arise due to multisensor misalignment, inadequate appearance forms, and perception-oriented issues, which complicate the corresponding relationship, information representation, and task-driven fusion. In this context, the advancement of artificial intelligence (AI) has driven the development of information fusion, offering a new perspective on tackling these challenges.1 AI leverages deep neural networks (DNNs) with gradient descent optimization to learn statistical regularities from multimodal data. By examining the entire process of multimodal information fusion, we can gain deeper insights into AI’s working mechanisms and enhance our understanding of AI perception in complex environments.


AI for misalignment information matching

While multimodal sensors provide multisource heterogeneous information for a complete scene representation, their limitations stem from factors like sensor position, angle, and inherent differences in sensor parameters. These factors lead to mismatches in different information, which adversely affect subsequent fusion. Fortunately, AI can learn statistical correspondences from existing multisource heterogeneous data, enabling adaptive matching in new scenarios.2 This data-driven learning approach indicates significant potential for robust information matching.


AI identifies correspondences in multisource heterogeneous data through two primary steps: feature extraction and correspondence regression (Figure 1A). These steps are closely interconnected, with effective feature extraction serving as a crucial prerequisite for accurate correspondence regression. However, the apparent heterogeneity in multisource data poses difficulties in extracting features that are conducive to regression. By discarding manually designed feature extractors, AI leverages DNNs to adaptively establish modality-invariant and intrinsic consistency presentations across different modalities. In terms of implementation, this can involve either explicit constraints or implicit implementations in an end-to-end manner. Naturally, regardless of the specific implementation, their ultimate objective is to extract modality-invariant and intrinsic consistency features. These features serve as the basis for subsequent regression of correspondences. Moreover, AI utilizing DNNs as regressors is a critical improvement. It frames correspondence finding as a function mapping problem, emphasizing a coarse-to-fine, multilevel approach. By integrating dense matching, semi-dense matching, or other parameter estimation methods, DNNs can regress discrepancies in multisource heterogeneous data, forming the basis for correction and subsequent fusion.




Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(7838) Cited by(0)

Relative Articles