Abstract:
As technology advances and human needs evolve, the demand for object detection across an expanding number of categories has significantly increased. However, fine-tuning models using only new data often results in catastrophic forgetting, where the model loses previously learned knowledge. While knowledge distillation has emerged as a promising approach to alleviate this issue, traditional feature distillation techniques primarily focus on shallow features, paying less attention to the rich semantic information embedded in deeper features. To address this, an attention-based multi-scale selective feature distillation method is proposed to distill both shallow and deep features, making comprehensive use of feature information at various scales. An attention-based selective module is incorporated during the distillation process to dynamically emphasize important features and selectively distill them. This approach enables the model to balance performance across both new and previous tasks and is further combined with classification and localization distillation. Extensive experiments on the MS COCO dataset have demonstrated the effectiveness of the attention-based multi-scale selective feature distillation method.