Visual object tracking: Progress, challenge, and future

Illustration of visual object tracking and summary of its progress, current challenges, and potential future directions
Visual object tracking aims to continuously localize the target object of interest in a video sequence. As one of the most fundamental problems in computer vision, visual object tracking has a long list of critical applications including video surveillance, autonomous driving, human-machine interaction, augmented reality, robotics, etc., in which the tracking system provides the capacity to report target positions in real time for subsequent visual analysis. In the past decades, visual object tracking has been extensively explored and has witnessed considerable progress, especially in deep-learning-based tracking. Despite this, robust tracking remains challenging due to many factors. To provide the community an overview, in this commentary, we will discuss visual tracking from different aspects. Specifically, we will first summarize the recent advancements achieved in visual tracking from the perspectives of algorithm and dataset. Then, we will analyze the challenges that the tracking community faces in developing practical tracking systems. Finally, we will discuss several promising directions for future research on visual tracking. It is worth noting that visual object tracking is a broad problem and consists of many specific topics. In this commentary, we focus on the most popular single-modality (i.e., RGB), bounding-box-based tracking. Figure 1 illustrates the task of visual tracking and the organization of this commentary.
